At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Infrastructure builds the systems engineers depend on to ship stable, scalable, and efficient services. We're hiring a Senior Technical Program Manager to run cross-functional programs across our infrastructure and data platform teams. This role blends program delivery with product sense: you'll own the roadmap for your area, set priorities, and act as the voice of the customer back into how we build. Responsibilities Run infrastructure programs end to end, from kickoff through delivery Own the roadmap for your platform area: shape the strategy, sequence the work, and make the prioritization calls Drive data platform migration and modernization work, coordinating across engineering, data, and platform teams to keep dependencies and timelines under control Be the voice of the customer: partner with engineering teams across Lyft, surface their pain points, and feed that back into priorities and roadmaps Define success metrics and adoption goals, gather and document customer requirements, and make sure what ships actually solves the problem Build feedback loops with customer teams and turn what you hear into concrete improvements Partner with engineering and infrastructure leads to build plans, call out risks early, and keep stakeholders aligned Own program health: track milestones, surface blockers before they slip, and keep decision-makers in the loop Use your technical background in distributed systems and data infrastructure to ask sharp questions and build plans the team believes in Share in the team's release oncall rotation Experience 5+ years in Technical Program Management or a TPM/PM hybrid role A background in software, data, or systems engineering, enough to go deep with engineers Experience owning a roadmap: setting strategy, prioritizing across competing demands, and defining what suc
Jobs in Canada
Infrastructure And Mlops Engineer in San Francisco
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current infrastructure and mlops engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Code Quality team sits within the Developer Platform organization and owns the systems that keep DoorDash's codebase healthy and secure as it scales: static analysis, quality gates, test frameworks, regression infrastructure, and tooling. Our job is to make sure the signals engineers rely on before shipping — test results, coverage, performance feedback etc — are fast and trustworthy. The decisions we make about tooling and standards directly shape how confidently and quickly engineering teams at DoorDash can ship to production. About the Role We're looking for Software Engineers to help build and maintain the systems that validate code quality across DoorDash's engineering org, treating our tooling as a critical product for the engineers who rely on it every day: static analysis and quality gates, test frameworks and regression infrastructure. You’ll design the tooling and automation that will help derive trustworthy quality signals, integrate them into the development lifecycle, and make it easy for engineers to execute reliable, repeatable workflows. You will collaborate across the engineering org, partnering directly with the teams who use what you build to understand the accuracy, reliability and performance of their functionality. You will report into the Engineering Manager on our Code Quality team in our Developer Platform organization. You must be located in either San Francisco, CA, Sunnyvale, CA, Los Angeles, CA, Seattle, WA, or New York, NY. You're excited about this opportunity because you will… Build and maintain quality tooling — static analysis, quality gates, coverage reporting, test frameworks, regression infrastructure — and integrate it directly into our developer workflows and CI/CD pipelines Define and derive quality signals - flakiness, pass rate, coverage, performance, scale readiness etc - Build tooling that improves everyday engineering workflows, including local development, CI/CD, debugging, and rollou
C$113.4K – C$162K/yr
We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste
$110K – $150K/yr
Location: San Francisco, CA (hybrid) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role We're looking for a highly analytical Senior Business Operations Analyst to support the growth and execution of our Dispatch Intelligence product. This role sits at the intersection of business operations, customer, product, engineering, and data science and has two core areas of responsibility. First, you will help drive overall program execution for Dispatch Intelligence: bringing structure to complex cross-functional initiatives, improving processes, tracking progress, and ensuring teams stay aligned on priorities and timelines. Second, you will help build and operate the processes through which flexible energy assets are onboarded onto the Verse platform and continuously improve their operational performance. The ideal candidate combines strong analytical problem-solving with exceptional project management and is comfortable working across both technical and commercial teams. You will play a critical role in helping Verse scale Dispatch Intelligence from individual projects and assets
About Prophecy Prophecy is building the next generation AI-powered data prep and analysis platform. Our platform enables business analysts and data teams to transform raw data into reliable, production-ready datasets and insights faster, using modern data infrastructure and AI-driven capabilities. We work with leading enterprises to simplify how organizations prepare, analyze, and operationalize data, while maintaining strong governance, security, and operational control. Our mission is to make it dramatically easier for organizations to turn complex data into trusted insights that drive decisions. About the Roles We are looking for a Director of Strategic Partnerships who can do both: drive revenue through Prophecy's partner ecosystem, and build an effective partner program. This is an early-stage, high-ownership motion, the playbook is still being written, and you'll have real influence over how we engage partners, what good looks like for partner-sourced pipeline, and how we build durable co-sell relationships with Snowflake, Databricks, GCP field and partner teams. You will sit at the intersection of sales, partnerships, and strategy, owning partner performance while building the programs and processes that scale it. You'll work directly with our AEs and SEs to bring partners into deals at the right moments, and you'll serve as the primary point of contact for our strategic cloud and ecosystem partners. What You’ll Own Partner Revenue & Pipeline Own and exceed partner-sourced and partner-influenced revenue targets on a quarterly basis Proactively generate pipeline through Snowflake, Databricks, and GCP field AEs and PDMs — building the relationships that produce qualified, sourced opportunities Drive joint account mapping and target account activation against Prophecy's ICP: enterprises running Alteryx on Snowflake or Databricks Activate co-sell motions through marketplace programs (GCP Marketplace, Snowflake Partner Network, Databricks
About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We are hiring a Firmware Validation & Integration Engineer for our autonomy software team. This is a critical role to build robust and scalable validation for our firmware and systems to ensure reliability at every level. In this role, you will work with our electrical, firmware, and autonomy engineers to build the infrastructure and test suites required to validate the system. This includes designing and implementing our Hardware-in-the-Loop (HIL) simulation environments and automation frameworks from the ground up. You will report to the Autonomy Platform Lead on our Autonomy Platform Team at DoorDash Labs. We expect this role to be hybrid with some time in-office and some time remote. You’re excited about this opportunity because you will… Play an integral role on a small and focused team. Design and build Hardware-in-the-Loop (HIL) systems to simulate vehicle dynamics and sensor data for comprehensive firmware and system-level validation. Develop automated test infrastructure and software tools to exercise multiple embedded platforms throughout our robot system. Interface many layers of our control system including vehicle controls, power management, and motion control to ensure seamless system integration. Implement low-level test sequences and validation algorithms to safely stress-test vehicle components such as batteries, drive-train, and thermal management devices. Collaborate with cross-functional teams to identify edge cases and hardware-software corner cases that impact vehicle safety and performance. We’re excited about you because… BS/MS degree in Computer Science, Robotics, Electrical Engineering, or related technical field. 5+ years of experience in validati
C$3 – C$7/hr
You will: Manage the growth program from end to end, including top-of-the-funnel growth / recruiting campaigns, applications processing, and hiring and onboarding Represent and champion the brand, promoting its value proposition to candidates and stakeholders, driving engagement, and taking ownership of process improvements that streamline growth pipelines and support program success Optimize full growth funnel for conversion and experience Build and lead new growth recruiting campaigns and channels to meet business goals Be the subject matter expert for all recruiting systems (ATS: Greenhouse), tools, and processes, and provide training/onboarding as needed Own internal reporting and analytics, and keeping all the internal hiring data clean and up-to-date in our systems Be the strategic driver behind on impactful initiatives to improve the workflow, data infrastructure and reporting Proactively flag discrepancies in hiring plans, interview process, JDs, offer details, etc. and maintain data integrity and operational excellence within the team Work from the San Francisco office, with occasional travel for onsite growth events. Ideally you’d have: Minimum of 3-7 years of experience working in Growth, Recruiting or RecOps at a rapidly growing company Extensive experience using Greenhouse as a Site Admin (user permissions, approvals, custom options) and pulling ad hoc reports using our internal TA tools (report connector, reporting capabilities, and limitations) Strong knowledge of Gsuite (VLOOKUP, pivot tables, data validation, conditional statements and formatting, filters, etc.) Excellent written and verbal communication skills, with the ability to tailor messaging to diverse audiences Must have a deep understanding of Growth / Recruiting Pipelines, Funnels and Conversion metrics Demonstrated excellent project management skills, ability to pull and manipulate data sets, critically analyze existing processes, and identify opportunities for process improvement Able to
From $250K/yr
About Scale Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust. About the Team Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities — and this role is not scoped to any single one of them. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like. About the Role As a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical needs — wherever the hardest ML problem in agentic AI happens to be that quarter. This could mean training and fine-tuning models, designing evaluation and observability systems, building improvement loops from production data, prototyping novel agent architectures, or designing internal systems and tooling that boost productivity across teams. You’re not tied to one team’s roadmap; you’re expected to move to where the technical leverage is highest, and t
From $216K/yr
The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai
About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and fi nancial reporting. Team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Sta ff Software Engineer,Data to be a technical lead and help architect and scale our data reliability, data infrastructure, automation and tools to meet growing business needs. You’re excited about this opportunity because you will... Own critical data systems that support multiple products/teams Develop, implement and enforce best practices for data infrastructure and automation Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Improve the reliability and scalability of our Ingestion, data processing, ETLs, Reporting tools and data ecosystem services Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We’re excited about you because... 8+ years of professional experience as a hands-on engineer and technical leader leading multiple projects 6+ years experience working in data platform and data engineering or a similar role You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Pro fi ciency in programming languages such as Python/Kotlin/Scala 4+ years of experience in ETL orchestration and work fl ow management tools like Air fl ow Expert in database fundamentals, SQL, data reliability practices and distributed computing 4+ years of experience with the Distributed data/similar ecosystem (Spark, Presto) and streaming technologies such as Kaa/Flink/Spark Streaming Excellent communication skills and experience working
Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Team DevX owns engineering velocity at Amplitude: build systems, CI/CD, developer environments, and internal tooling. We're building a software factory: automated workflows that remove manual bottlenecks from how engineers ship code. The Role We're looking for a Staff Software Engineer – DevX (Hybrid – San Francisco) who bridges infrastructure and application thinking and can accelerate how the whole team develops, tests, and ships in the cloud. You'll set architecture for our developer platform, lead our software factory work, and push our development model toward cloud-first workflows. This is a high-leverage, low-oversight role. You'll own initiatives end to end, from an ambiguous problem to production, and set technical direction for a foundational team. What You'll Do Cloud development platform: Unify and scale our existing loca
$150K – $240K/yr
Location: San Francisco, CA (Remote/Hybrid Available) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of the brightest industry experts in the field building cloud-native applications that scale to trillions of data points collected from electricity markets globally. You will be a part of a dynamic, robust team primarily supporting the backend needs of our Aria software product spanning hundreds of data sources, sinks, services, and jobs. Your expertise will not only have a direct impact on product decisions, but you also be well-positioned to drive the development and trajectory of our entire platform and infrastructure and influence important architectural decisions that affect the whole organization. Key Responsibilities Foster a culture and mindset of well-designed systems, test-driven software, and transparent communication with a high caliber of mutual respect and consideration for stakeholders Read and write a lot of Go, Python, and Protobuf Build, test, debug, maint
$150K – $240K/yr
Location: San Francisco, CA (Remote/Hybrid Available) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role You will be a member of the technical staff developing product experiences for our Dispatch Intelligence users – customers who want and have battery energy storage systems for additional energy cost savings or faster interconnection times. In this role, you will serve in a “full stack” capacity designing, building, and maintaining frontend and backend components of our energy storage suite of applications. We use Typescript, React, Next.js, Tailwind CSS, Radix/ShadCN, Jest, Cypress, Playwright, Vitest, and Storybook with Echarts and D3/Observable for data visualization for our frontend, and Cloudflare Pages for hosting and content delivery. We rely on identity and auth platforms like Clerk for sign-in flows. Our backend is written in Go and Python with Postgres/AlloyDB and blob storage for data persistence. Key Responsibilities Foster a culture and mindset of well-designed systems, test-driven software, and proactive communication with a high degree of transparency, mutual resp
$150K – $210K/yr
Location: San Francisco, CA (Hybrid) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role We're seeking an experienced Senior Optimization Engineer to join our Data Science Team. In this role, you will lead the design, development, and deployment of optimization models that power our software platform across applications including electricity markets, renewable energy, and battery energy storage systems. You will be responsible for developing production-grade optimization engines that solve complex operational and planning problems at scale. This role requires deep expertise in mathematical optimization, strong software engineering skills in Python, and experience building optimization models that integrate with production systems. The ideal candidate has significant experience in the energy industry, particularly electricity markets and battery storage optimization. This position emphasizes technical leadership, ownership of complex optimization projects, and collaboration across engineering, product, and commercial teams to deliver high-impact optimization solutions. Key Res
$150K – $240K/yr
Location: San Francisco, CA (Remote/Hybrid Available) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focused on Fleet Telemetry & Control at Verse, you will be working closely with our energy solutions partners to design, implement, and test distributed energy resource controls and telemetry software on customer hardware at sites around the world. You will be part of a dynamic, high-performance team building applications directly on bare-metal or on hardware-level virtualization platforms. As an advanced technical leader in network programming and state management development, engineering teams will look to you for best standards and practices for interfacing with on-premises grid assets using solutions you will build and maintain. Key Responsibilities Foster a culture and mindset of well-designed systems, test-driven software, and transparent communication with a high caliber of mutual respect and consideration for stakeholders Mentor and support career and junior level engineers in their fleet telemetry and control software career development
Other cities to consider
More places hiring for this role
Get new infrastructure and mlops engineer jobs in San Francisco, Canada by email
Daily job updates · Unsubscribe anytime