About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
Jobs in Canada
Infrastructure Team Manager in San Francisco
128 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure team manager jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
From $179.4K/yr
Scale GP (Scale Generative AI Platform) is an enterprise-grade AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our core infrastructure in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in diverse, highly-regulated industries like healthcare, telecom, finance, and retail. What You’ll Do: Architect multi-cloud systems and abstractions to allow the SGP platform to run on top of existing Cloud providers Implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Collaborate with platform, product teams and our customers directly to develop and implement innovative infrastructure that scales to meet evolving needs. Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Be able, and willing, to multi-task and learn new technologies quickly What We’re Looking For: 4+ years of full-time engineering experience, post-graduation Experience scaling products at hyper growth startups Experience tinkering with or productizing LLMs, vector databases, and the other latest AI technologies Proficient in Python or Javascript/Typescript, and SQL Experience with Kubernetes Experience with major cloud providers (AWS, Azure, GCP) Excellent communication skills with the ability to explain technical concepts to both technical and non-technical audiences Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries fo
From $184K/yr
Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our dynamic Public Sector Engineering team. As a part of this team, you will own the model inference layer - enabling state of the art models, debugging the latest AI tools, managing networking, debugging latency, and tracking pricing/usage metrics for AI models. You will lead technical discussions on the frontlines with cloud vendors and customers to deliver on critical contracts and to debug platform issues. You will also work upstream with Product to understand features before they break, moving us from "infra-only debugging" to proactive integration testing. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. You will work with Product to build integration tests that catch issues early, shifting the focus from "infra-only debugging" to preventing failures upstream. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Proficiency in both front-end and back-end development, including experience with modern web develo
$140K – $225K/yr
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust
About the Team Come help us build and develop tools serving hundreds of engineers internally! We’re looking for a Fullstack Software Engineer to join our Developer Insights team. About the Role Our mission is to improve the developer experience of engineers at DoorDash by building various internal products, including our internal developer portal, Developer Insights. Our success as a platform team depends on the success of the product teams we serve. Because of this, we invest in building a strong community that encourages participation and promotes best practices. You’re excited about this opportunity because you will… Introduce cutting edge technologies to our engineering organization, including tools built on LLMs Build new features for Developer Insights (using Backstage.io) Improve the developer experience for all of our engineers Work and collaborate across team boundaries. Contribute features and bug fixes to upstream open-source projects. Mentor and educate your peers. Lead the team in a technical fashion and assist in roadmap planning and measurement of existing features. Represent the team at large in OKR and engineering all-hands presentations. Context switch from frontend to backend to data depending on the need that arises. We’re excited about you because… You have at least 2 years of experience in web technologies using Typescript with React on the frontend with Java, Kotlin, Python or Go backend experience. You have a product mindset and apply that to how you would build out platform services. You love systems and software, and you're proficient in both. You’re curious and dive deep into different system architectures. You are an organized and excellent written and verbal communicator. You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Compensation The successful candidate’s starti
$165K – $247K/yr
Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Resp
About the Team DoorDash is building the world’s most reliable on-demand logistics engine for delivery! We’re looking for machine learning engineer interns to join our fast-growing engineering team to help us develop a 24x7 global infrastructure system that powers DoorDash’s three-sided marketplace of consumers, merchants, and dashers. About the Role As a Machine Learning Engineer intern at Doordash, you’ll work on tackling new challenges in machine learning and artificial intelligence. You’ll conduct research that can be applied across Doordash engineering teams and engage in external collaborations and mentoring, while also performing research in any of the following areas: Auction, Game theory, Recommender systems, Ranking, AdTech, Computer Vision, Causal Inference, and Big data analytics. We offer a 12-week summer internship program in our San Francisco, Sunnyvale, New York, or Seattle offices. You’re excited about this opportunity because you will… Use cutting-edge research in ML/AI, NLP, RecSys, Ranking, Computer Vision, Causal Inference, Ad Tech, Graph analysis to solve real-world problems across discovery, ads,forecasting, fulfillment and search experiences at Doordash. Contribute and execute on research ideas that can be applied and used to improve product experience at Doordash. Collect, analyze, and synthesize findings from data and use these insights to build relevant ML models. Write clean, efficient, and sustainable code We’re excited about you because you… Are working towards a Masters degree in Computer Science, ML, NLP, Statistics, Information Sciences or related field and are graduating between Fall 2027 & Summer 2028 Have a mastery of at least one systems languages (Java, C++, Python, Kotlin, GoLang) or one ML framework (Tensorflow, Pytorch, MLFlow) Have experience in research and in solving analytical problems Are a strong communicator and team player. Have a passion for applied ML and the Doordash product Ideally have
About the Team DoorDash is building the world’s most reliable on-demand logistics engine for delivery! We’re looking for machine learning engineer interns to join our fast-growing engineering team to help us develop a 24x7 global infrastructure system that powers DoorDash’s three-sided marketplace of consumers, merchants, and dashers. About the Role As a Machine Learning Engineer intern at Doordash, you’ll work on tackling new challenges in machine learning and artificial intelligence. You’ll conduct research that can be applied across Doordash engineering teams and engage in external collaborations and mentoring, while also performing research in any of the following areas: Auction, Game theory, Recommender systems, Ranking, AdTech, Computer Vision, Causal Inference, and Big data analytics. We offer a 12-week summer internship program in our San Francisco, Sunnyvale, New York, or Seattle offices. You’re excited about this opportunity because you will… Use cutting-edge research in ML/AI, NLP, RecSys, Ranking, Computer Vision, Causal Inference, Ad Tech, Graph analysis to solve real-world problems across discovery, ads,forecasting, fulfillment and search experiences at Doordash. Contribute and execute on research ideas that can be applied and used to improve product experience at Doordash. Collect, analyze, and synthesize findings from data and use these insights to build relevant ML models. Write clean, efficient, and sustainable code We’re excited about you because you… Are working towards a PhD degree in Computer Science, ML, NLP, Statistics, Information Sciences or related field and are graduating between Fall 2027 & Summer 2028 Have a mastery of at least one systems languages (Java, C++, Python, Kotlin, GoLang) or one ML framework (Tensorflow, Pytorch, MLFlow) Have experience in research and in solving analytical problems Are a strong communicator and team player. Have a passion for applied ML and the Doordash product Ideally have work
About the Team DoorDash is a data driven organization and relies on timely, accurate and reliable data to drive many business and product decisions. The Core Data Platform organization owns all the infrastructure necessary to run an operationally efficient analytical data stack. About the Roles The Data Platform team spans data mobility frameworks, ingestion, infrastructure, tools, and governance. Together, they design and operate scalable compute and ingestion frameworks using technologies such as Spark, Flink, Kafka, Airflow, and modern lakehouse solutions, while also building abstractions and tools that simplify data workflows for engineers, analysts, and ML practitioners. In parallel, these teams establish strong data quality, cataloging, privacy, and compliance standards to ensure trust in analytics and regulatory adherence. As relatively high-impact teams, they offer engineers the opportunity to shape the roadmap, influence core platform decisions, and directly enable DoorDash’s business-critical insights and real-time personalization capabilities. You must be located in San Francisco, CA, Sunnyvale, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Drive vision & strategy for building the frameworks charter and position it to handle the challenges of a rapidly growing business. Scale the analytical platform for the increasing amounts of data and use cases. You will bring your expertise in building and operating high scale systems with a focus on reliability, scalability and cost efficiency. Collaborate with stakeholders building solutions on top of the platform Foster a positive and supportive work culture, upleveling others. We're excited about you because you have… B.S., M.S., or PhD. in Computer Science or equivalent. 2+ years of industry experience at our I4 level, 5+ years of industry experience at our I5 level Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in th
About the Team DoorDash Labs is an independent team within DoorDash. We're hiring a backend software engineer to work at the intersection of software engineering and robotics to solve key business problems with elegant technical solutions. If you have a passion for applying robotics solutions to a service loved by millions of people, then we want to talk to you! About the Role We’re looking for Backend Engineers to work on both Product and Product Platform based teams in DoorDash Labs. Product focused Engineers work at the intersection of product and infrastructure to solve key business problems with elegant technical solutions. You'll operate our backend services and architecture that support all product functionality and will be challenged to consider the big picture -- collaborating cross-functionally, as well as evaluating and executing on trade-offs to maximize business impact for the company. You're excited about this opportunity because you will... Design and implement backend services for IoT that integrates with core DoorDash data, focused on reliability, and future extensibility Create a well documented APIs for other departments to integrate with Improve performance, reliability, scalability and security for our backend systems Introduce tools and best practices to accelerate our development process Design and implement backend services for autonomous delivery system that integrate with core DoorDash data. We're excited about you because you have... B.S., M.S., or PhD. in Computer Science or equivalent 6+ years of industry experience as a software engineer Experience with backend for frontend architecture Ability to improve efficiency, scalability, and stability of multiple system resources Experience with service oriented architecture, writing REST API’s, unit testing, and architectural design Understanding of modern web stacks and architecture (HTTP, REST) Experience with SQL Experience with either Java or Kotlin Nice to Have Experience with
From $102K/yr
About the Team The DoorDash Research Fellowship is a 3-month program (extendable to 6 months) looking for Summer and Fall 2026 cohorts, for researchers and engineers who want to work on the hardest applied ML and AI problems in local commerce. Fellows are given the resources, autonomy, and access to real-world operational data needed to pursue ambitious research directions — with the goal of producing work that influences both the field and how DoorDash operates at scale. This program is modeled on the best external research fellowships: fellows are treated as independent researchers, not as junior employees on a product team. You pick the problem (within a set of priority areas), you own the direction, and you publish or ship the outcome. You’re excited about this opportunity because you will receive… Dedicated compute allocation sized to the research agenda — GPU clusters for training and inference budgets for experimentation Full access to DoorDash's research infrastructure — our internal RL stack, training and evaluation pipelines, RL environments built on real operational systems, agent evaluation harnesses, and the tooling our own research teams use day-to-day. Fellows are first-class users, not sandboxed visitors. Access to DoorDash operational data — real-world datasets spanning logistics, merchant operations, consumer behavior, and marketplace dynamics, under appropriate data governance Research mentorship from senior researchers and engineering leaders at DoorDash, plus a named research sponsor for each fellow who meets with you weekly and is accountable for unblocking your work Speaker series featuring leading researchers and practitioners from academia and industry — faculty from top ML programs, research leads from frontier AI labs, and senior operators from across tech. Fellows get dedicated 1:1 time with speakers when possible. A cohort of fellows working alongside you — a small, tight-knit group of researchers tackling different problems but sharing
About the Team The Consumer Engineering Team is responsible for helping consumers discover and order everything they love globally. Our work spans the entire consumer journey across homepage, search, store discovery, item exploration, checkout and post checkout. We aim to craft a hyper-personalized, delightful and frictionless experience for millions of our customers. About the Role As a Senior Staff Machine Learning Engineer on Core Cx, you will set the personalization (P13n) strategy for the entire consumer shopping journey and bring that strategy to life. You will use our robust data and machine learning infrastructure to implement new ML solutions to make the consumer search experience more relevant, seamless, and delightful across restaurant, grocery, retail and all business at DoorDash . You will modernize the recommendation system leveraging AI. You will demonstrate a strong command of production level machine learning, experience with solving end-user problems, and collaborate well with multi-disciplinary teams. You're excited about this opportunity because you will… Drive the engineering vision, strategy, and execution for an organization of 150+ Grow, build, and nurture impactful business-focused product engineering teams. Scale the team by developing leaders internally and attracting world-class talent Mentor and guide a fast-growing organization in setting the right architectural patterns, working with various vendors in the space, and making judicious investments in the right areas anticipating what the company needs a few years down the road. Partner with Business, Product, and other Engineering teams to transform DoorDash from local commerce to agentic commerce We're excited about you because you have… B.S. or M.S. in Computer Science or equivalent. 10+ years of industry experience developing machine learning models with business impact, and shipping ML solutions to production. Proficiency in using AI coding tools (e.g., Claude Code) in th
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The AI team is focused on one of Stripe’s most strategic growth areas: enabling the monetization and scaling of AI-native and AI-enabled businesses. We’re in a unique position – partnering with the world’s most ambitious AI companies (the likes of OpenAI, Anthropic, NVIDIA, etc) building on the frontier of artificial intelligence -- across infrastructure, foundation models, agents, and applications -- to help them grow and commercialize globally using Stripe’s full financial stack. As part of Stripe’s GTM / Sales organization, this team works closely with Product, Engineering, and Marketing to shape Stripe’s AI GTM strategy and ensure that the world’s leading AI companies -- from early-stage innovators to the largest public players -- choose Stripe as their monetization platform. What you’ll do Work with existing Stripe customers in the AI Industry to develop and execute long-term sales strategies to expand Stripe’s revenue Own the full sales cycle, from business case development, to deal structuring and negotiating, to close Develop account plans and cross sell into your list of strategic AI customers, driving growth through expansion and new revenue streams Drive deal strategy and commercial negotiations for large, complex renewals Develop relationships with executive stakeholders within your book of business, deeply understanding problems they are solving and helping drive to solutions Be responsible for account mapping and coordinating ef
From $180K/yr
Scale GP is Scale's enterprise Generative AI platform—APIs and infrastructure for knowledge retrieval, inference, evaluation, and intelligent automation. We power mission-critical workflows for leading enterprises, helping teams turn complex data and models into reliable, production-ready AI systems. We're building a new AI Enablement team to create the next generation of agent-powered tools that ground AI in real operational workflows. Our goal: help internal teams demystify their own workflows, then deploy agentic systems that reason over data, take action, and deliver measurable outcomes. We don't build in a vacuum. You'll use our own platform to solve real business problems internally—then selectively commercialize that same stack for customers. What we run on is what we sell. This is a 0→1 team. We're looking for a sharp, product-minded engineer who thrives in ambiguity, moves fast, and loves building systems from scratch alongside customers and cross-functional partners. You'll work closely with product, forward-deployed engineers, data scientists, and applied AI teams to turn real-world problems into scalable production solutions. If you like shipping fast, owning outcomes, and working across the stack—from polished frontends to distributed backends to LLM integrations—this role is for you. What You’ll Do Own full-stack features and projects end-to-end — from design through production deployment — within a larger product area Sample surfaces - Accounting Agents, Finance Copilots, GTM Agents, Agentic Experimentation Platforms Develop reliable backend services in Typescript/Python, work with distributed systems, data pipelines, and AI/ML infrastructure Integrate LLMs, vector databases, and agentic frameworks to power intelligent workflows Ship quickly through tight experimentation loops while maintaining high quality and reliability Adapt across the stack and learn new tools as needed to solve real problems end-to-end Ideal Experience 3+ years of full-tim
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The AI team is focused on one of Stripe’s most strategic growth areas: enabling the monetization and scaling of AI-native and AI-enabled businesses. We’re in a unique position – partnering with the world’s most ambitious AI companies (the likes of OpenAI, Anthropic, NVIDIA, etc) building on the frontier of artificial intelligence -- across infrastructure, foundation models, agents, and applications -- to help them grow and commercialize globally using Stripe’s full financial stack. As part of Stripe’s GTM / Sales organization, this team works closely with Product, Engineering, and Marketing to shape Stripe’s AI GTM strategy and ensure that the world’s leading AI companies -- from early-stage innovators to the largest public players -- choose Stripe as their monetization platform. What you’ll do Work with existing Stripe customers in the AI Industry to develop and execute long-term sales strategies to expand Stripe’s revenue Own the full sales cycle, from business case development, to deal structuring and negotiating, to close Develop account plans and cross sell into your list of strategic AI customers, driving growth through expansion and new revenue streams Drive deal strategy and commercial negotiations for large, complex renewals Develop relationships with executive stakeholders within your book of business, deeply understanding problems they are solving and helping drive to solutions Be responsible for account mapping and coor
Other cities to consider
More places hiring for this role
Get new infrastructure team manager jobs in San Francisco, Canada by email
Daily job updates · Unsubscribe anytime