ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the first people who will help define it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for building great AI Infrastructure. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. The role Getting a model into production still takes real expertise — choosing a serving engine, sizing hardware, tuning it, wiring it into an app. We want a developer to go from "it runs on my laptop" to "it's serving production traffic" in minutes, on their own. You'll own the entire experience a developer touches to deploy and iterate: the CLI and SDKs, the console, onboarding, model discovery, deployment configuration, truss, and the increasingly agent-driven ways developers build. Your job is to make Baseten synonymous with Great DevEx and make it effortless to drive and self-serve deploy models on Baseten for far more developers than it is today. Impact and outcomes you'll drive You will collapse time-to-production — take a developer from first sign-up to a running, maint
Jobs in United States
Deployment Lead in United States
636 active opportunities · Updated October 2026
Showing
15 jobs
Explore current deployment lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. PRODUCT AT BASETEN Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and infrastructure engineers in the world, and you'll set the standard for what product looks like here. You earn trust by being technical, finding the truth in front of customers, building great cross-functional relationships, and shipping great product experiences. THE ROLE The largest, most demanding Enterprises are starting to run on Baseten and they come with a range of security, compliance, and procurement requirements. Today that readiness is assembled deal-by-deal. You'll own the enterprise-readiness surface end to end and turn it into product: the deployment options customers can choose and buy, compliance posture they can trust, access and security controls their IT teams require, and the billing and spend controls their finance teams expect. What does a complete Baseten Enterprise Product offering look like? RESPONSIBILITIES Drive the Enterprise Readiness customer experience end to end: Partner with GTM and Enterprise Engineering to make "enterprise-ready" a platform-wide capability, not a deal-by-deal scramble. Outcome: readiness becomes a supported, priced product instead of bespoke work asse
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As an early member of Baseten's Platform Team, you will be pivotal in building internal infrastructure to support our engineering organization. You will own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across Baseten to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. If you are passionate about elegant solutions—like streamlined monorepos, lightning-fast CI pipelines, and thoughtfully designed shared libraries—you'll thrive at Baseten. RESPONSIBILITIES Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 This role bridges infrastructure and product engineering: you'll build genuine partnerships across product, AI, and enterprise teams so that ownership is shared and velocity is never blocked by platform constraints. As Director, you'll set a forward-looking cloud vision, proactively align with stakeholders across the business, and ensure the platform scales for multi-shard, multi-region growth while meeting security and compliance commitments (SOC 2, business continuity/disaster recovery, and enterprise security frameworks). The Role: Cost Efficiency: Drive significant annual infrastructure savings by migrating data workloads to EKS and self-hosting key services such as OpenSearch. Ingress Convergence: Deprecate legacy frontend ALBs and consolidate to a single EKS-managed ALB per shard, unblocking faster deployments across the org. Coverage & Bench Depth: Eliminate single points of ownership across Networking, OpenSearch, and Terraform through cross-training and targeted hiring into coverage gaps. Stakeholder Alignment: Stand up a recurring alignment cadence with Product, AI, Enterprise, and Security so infra planning is driven by demand, not ad-hoc interrupts. Automation: Ship automated shard buildout via Backstage to remove manual toil from enterprise scaling. Reliability: Cut P0/P1 incidents attributed to Cloud Platform (DNS, ALB misconfiguration) through hardened ingress patterns and Terraform-policy guardrails, including blocking unauthenticated public endpoints. Roadmap Ownership: Deliver a roadmap covering Agent enablement, centralized IaC, and deployment rollout acceleration, tied to AI and
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Job Summary We are looking for a GTM DevOps Engineer to join our Business Systems team and own the reliability, automation, and delivery infrastructure behind our Go-To-Market (GTM) technology stack. This role sits at the intersection of platform reliability and CI/CD engineering, ensuring that our critical business systems — including Salesforce, NetSuite, MuleSoft, Workato, and an expanding portfolio of AI-powered workloads — are deployed consistently, operate resiliently, and scale with the business. You will partner closely with Business Systems developers, architects, and business stakeholders to build and maintain the pipelines, monitoring frameworks, and operational standards that keep our GTM systems healthy and our release cycles fast and predictable. As our team builds and deploys AI agents across GCP Cloud Run and AWS Bedrock AgentCore, you will serve as the infrastructure and deployment owner for these workloads — bringing engineering discipline to an environment where AI-generated code is increasingly entering production. This is a hands-on engineering role for someone who thrives in complexity, takes ownership of platform uptime, and brings a software engineering mindset to business application operations — directly supporting GTMSOE's broader mission of operational excellence across the GTM org. Key Responsibilities CI/CD & Release Engineering Design, build, and maintain CI/CD pipelines for Salesforce (SFDX/Salesforce CLI), NetSuite (SuiteScript/SuiteBundler), MuleSoft (Anypoint Platform), and Workato; establish branching strategies, environment promotion standards, and release gatin
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
From $10K/yr
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The team owns the core experience that helps finance teams control, automate, and optimize company spend. We build the engine that powers Ramp’s card and expense workflows—from card issuance and spend limits to approvals, policy enforcement, and real-time insights. Our systems handle billions of dollars in transactions and integrate deeply with Ramp’s AI platform, financial infrastructure, and partner ecosystems (banks, ERPs, HRIS). As a backend engineer on this team, you’ll work across product surfaces that are central to Ramp’s success: cards, approvals, spend controls, and automation intelligence. What You’ll Do Design, build, and scale backend systems that power spend controls, approval workflows, and card transactions at massive scale Collaborate cross-functionally with product, design, and data to deliver intelligent, user-first experiences for finance teams Integrate with Ramp’s internal AI platform to automate spend policy enforcement and anomaly detection Own complex projects end-to-end — from architecture to deployment and o
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. The Mandate: Industry Disruption and 10x Scale: Vision: Product Foundry is Replit's internal engine for innovation . We move Replit beyond being a collaborative development environment to becoming the foundational operating system for the entire next generation of software, where AI Agents are the primary actors. The 10x Goal: Our purpose is to launch high-risk, full-stack 0→1 initiatives that define and establish entirely new, multi-billion-dollar product categories. We are seeking non-linear growth opportunities, fundamentally aiming to 10x Replit's value and addressable market by proving out unprecedented technical primitives and disruptive Go-To-Market strategies. The Audience: We build for the next generation of creators and high-leverage users and enterprises, equipping them with tools that enable them to build anything, anywhere . Candidates that do well here will certainly go on to build their own companies in the future! This is a high visibility role reporting to Execution Model: High-Agency Founding Teams Structure: We operate as a collective of in-house technical founders —not just specialized engineers. Initiatives are run by lean, autonomous squads built for velocity and maximum technical leverage. This model is centered around an Engineer DRI (Directly Responsible Individual) who maintains total ownership over the initiative's technical, product, and launch success, supported by fractional PM and Design resources. Cadence: We enforce rapid iteration and rapid market validation via 3-week sprints per initiative. This cadence forces fast deployment, immediate user feedback, and tight alignment with the internal betting table process, mirroring the intensity and speed of a lean startup. Required skills and
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About this role: Accelerate engineering velocity and reduce friction in the Replit development experience by stewarding our codebases, tooling, and developer workflows. You will directly impact every engineer who ships code at Replit. By focusing on developer experience, you will act as a force multiplier across all product development teams, enabling faster feature delivery, happier developers, and reduced operational overhead. The role combines the technical depth needed to navigate a complex, polyglot codebase with the product mindset to understand how infrastructure decisions impact developer productivity and ultimately customer value delivery. You will also partner closely with the AI team on our internal AI platform — which already generates more than 60% of all merged PRs at Replit — to improve the Agent's output and help shape strategy around the Agent's default stack. This is an early hire in this area, so you will have agency and have an accelerated career path as the team undoubtedly grows. You will: Maintain and evolve our codebase structure — a complex TypeScript monorepo, Go services, npm packages, and internal Agentic tooling. Own the build and test pipelines and optimize them to minimize build times and improve developer iteration speed. Drive code generation and type-safe interfaces across service boundaries (e.g., GraphQL, Protocol Buffers, gRPC, OpenAPI). Set the standards for code quality using automation tools such as TypeScript, ESLint, Prettier, and Go linters/formatters — building custom rules and plugins to enforce Replit-specific requirements. Streamline development setup and the onboarding experience. Work with platform teams to improve deployment processes, infrastructure integrations, and e
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Job Summary We are looking for an experienced Growth Infrastructure Engineer to build and maintain the technical backbone that enables scalable growth experiments, high-performance data pipelines, and automated systems that drive user acquisition, engagement, and product iteration. This role sits at the intersection of growth, product, and infrastructure — combining deep technical engineering with experimentation and data-driven optimization. You will collaborate with product, data science, and backend teams to ensure that growth initiatives run smoothly and scale efficiently across systems. Key Responsibilities Growth Infrastructure & Systems Design, implement, and maintain scalable infrastructure that supports growth and experimentation needs. Build and optimize analytics pipelines to capture key product and growth metrics (acquisition, activation, retention, etc.). Develop automated workflows for user onboarding, campaign delivery, and performance tracking. Experimentation & Optimization Support A/B testing frameworks and integrate them into production systems. Enable reliable data collection and evaluation for growth experiments. Automate deployment and rollout of growth feature flags and tests. Cross-Functional Collaboration Partner with Growth Product Managers, Data Engineers, and Analysts to define technical requirements for growth initiatives. Translate business goals into technical specifications and system designs. Provide guidance on performance, reliability, and scalability trade-offs. Monitoring & Reliability Implement monitoring and alerting for growth infrastructure services. Troubleshoot production issues and optimize for uptime and performance. Ensure data quality and consistency for report
What you’ll do Partner with medical image reconstruction scientists / engineers to build ML components that improve reconstruction quality, speed, robustness, or quantitative accuracy. Define training/evaluation pipelines, datasets, and metrics that map to user needs and design requirements. Productionize models: inference performance, reproducibility, monitoring for drift/regressions, and safe fallbacks. Collaborate on hybrid algorithms, incorporating physics and learned priors, denoisers, learned regularizers, and quality estimation. Help build tooling for rapid experimentation as well as rigorous verification of algorithm changes. What we’re looking for Strong applied ML experience plus comfort with signal processing / imaging or adjacent domains. Ability to move fluidly between research prototypes and production-quality systems. Strong evaluation discipline: metrics, ablations, data leakage avoidance, and reproducibility. A demonstrated track record of applying ML to physics-based or inverse problems (i.e., shipped projects, a portfolio, or publications.) Useful experience ML for imaging/inverse problems (or adjacent) with strong evaluation discipline and comfort with GPU performance constraints. Pragmatic production mindset: reproducible training/inference, regression testing, and safe deployment in high-stakes contexts. A background in computational physics or scientific computing. Leverage ML-based methods such as PiNNs and Neural Operators to solve partial differential equations arising in ultrasound simulation and imaging. Experience in Agentic-SciML is a plus. Hands-on experience with data curation for ML: building datasets from messy, real-world sources, defining ground truth, and managing labeling or simulation pipelines. Background in data assimilation: combining observations with physics-based models (Kalman filtering, variational methods, ensemble approaches, or learned variants).
From $88.1K/yr
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Platform Engineer, you will contribute to building and improving Stitch Fix’s cloud-native infrastructure and internal developer tooling. You’ll work on tools and automation that help product engineers deploy, operate, and debug services more easily, while learning modern platform engineering practices alongside experienced teammates. This role is ideal for engineers who enjoy improving developer experience and want to grow their skills in cloud infrastructure and CI/CD systems. Responsibilities: Contribute to the development and evolution of our internal platform-as-a-service used by application and service developers Build and maintain tooling that improves developer workflows, deployment reliability, and day-to-day productivity Collaborate with platform and application engineers to identify friction points and implement incremental improvements Learn and apply best practices around Infrastructure-as-Code, containerized workloads, and CI/CD pipelines Use, or are eager to adopt, AI-assisted development tools to improve productivity, and are excited to help explore and integrate LLM-powered solutions that automate internal support and operational workflows Have opportunities to propose ideas and improvements, with support and mentorship from the team Things you’ll get exposure to (and we don’t expect experience with everything): AWS Terraform, Pulumi CircleCI Docker, ECS, EKS Ruby, Golang, Python About You 2+ years of software development and infrastructure experience with significant contribut
From $399.4K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The platforms in our Engineering Acceleration org dictate how thousands of Roblox engineers ship, and how every backend service gets built, deployed, and kept reliable at scale. They sit on the critical path for production reliability, developer velocity, and security posture. Reporting to Andrew Swerdlow, you will rethink the entire engineering toolchain from the ground up to be agentic-native, creating a world where AI agents are first-class participants in software development and humans set direction and supervise. This is a unique opportunity to define what modern engineering infrastructure looks like at one of the largest platforms in the world. You will: Reimagine the engineering toolchain as agentic-native, designing the platforms, guardrails, and feedback loops that let AI agents safely drive migrations, validation, and routine operational work. Set a bold technical direction for AI-driven quality, including agent-generated tests, automated coverage of untested paths, intelligent verification, and continuous-deployment workflows. Make software quality and SEV prevention a measurable property of the platform by investing in safe-change mechanisms, automated verification, progressive
Other cities to consider
More places hiring for this role
Get new deployment lead jobs in United States by email
Daily job updates · Unsubscribe anytime