Jobs in United States

Ai Deployment Engineer in United States

5,082 active opportunities · Updated October 2026

Explore current ai deployment engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

Hiring demand

25/100

cooling · 8 related jobs

Hiring trend

-40%

Job postings compared with the previous 30 days

Remote options

37.5%

Share of matching jobs listed as remote

B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr

KubernetesRestMachine LearningAI
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

We are looking for a Senior Forward Deployed Engineer to join the Customer Solutions team. You will be the technical authority embedded with our most complex customers, guiding them through deployment, architecture, onboarding, and the adoption of agentic development workflows. You bring deep hands-on experience from prior roles and use that depth to advise, design repeatable patterns, and drive customer outcomes end to end. You operate autonomously, own the technical success of your customers, and bring their experience back to shape how Coder builds and delivers. This is not an execution-only role. You are equally comfortable doing deep technical work with a customer and stepping back to design the repeatable pattern behind it. You are energized by ambiguity, motivated by customer outcomes, and capable of influencing organizational change alongside the technical work. This position is required to sit in the Eastern Time Zone. What You'll Do Serve as the primary technical authority for post-sales customers, guiding deployment architecture, environment design, and adoption of Coder across both human and AI development workflows Own onboarding engagements end to end, ensuring customers move from contract to productive adoption with speed and confidence Lead Get Well engagements where architecture decisions, rollout patterns, or organizational dynamics are limiting customer health or growth Help customers implement the technical and organizational changes required to adopt agentic development practices at scale Design and document repeatable delivery patterns across onboarding, architecture, and adoption that can scale across customer segments Design and recommend reference architectures tailored to each customer's cloud environment, security posture, and organizational constraints Translate customer environment complexity into clear guidance on networking, ingress, identity, and infrastructure patterns Anticipate technical and operational risks, escalate to the right

AWSAzureKubernetesCI/CD
V
📍 United States· Full-time
✓ Quality checkedCompany trend -90%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta’s new Enterprise Resilience team is being formed to support the next era of growth by powering services that are reliable, scalable, and resilient by design. As our customer base expands and our systems scale, we need a dedicated group focused on partnering closely with product engineering teams to build and operate robust distributed systems across all of Vanta’s environments, including our new FedRAMP deployment. In this role, you’ll help define the foundations of reliability at Vanta including shaping best practices, building core infrastructure, and guiding teams as they design services that perform consistently for customers. This team will have a broad and deep impact across product engineering. Your work will influence how every Vanta engineer builds, deploys, monitors, and maintains their services, whether for our commercial environment or regulated customers with more stringent requirements. You’ll develop tools and frameworks that make it easier to detect and remediate issues, improve operational readiness, and support feature development that meets the needs of increasingly large and complex enterprise customers. Vanta engineers design and develop new product functionality and infrastructure using modern frameworks and tooling, including TypeScript, React, Node.js, MongoDB, GitHub Actions, and AWS services such as Fargate and ECS. If you're excited to help define a new function, raise the reliability bar across an entire engineering organization, and build systems that scale with Vanta’s growth, we’d love to meet you. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll

TypeScriptReactNode.jsMongoDB
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why this role? This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. We’re looking for Software Engineers with Applied AI experience who can own the design, build, and deployment of agentic workflows powered by Large Language Models (LLMs), from early prototypes to production-grade AI agents, to deliver concrete business value in enterprise workflows. You’ll work closely with customers on real-world business problems, often building first-of-thei

PythonReactGitAI
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. You may

PythonGitRestMachine Learning
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why this role? This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. We’re looking for Software Engineers with Applied AI experience who can own the design, build, and deployment of agentic workflows powered by Large Language Models (LLMs), from early prototypes to production-grade AI agents, to deliver concrete business value in enterprise workflows. You’ll work closely with customers on real-world business problems, often building first-of-thei

PythonReactGitAI
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Job Summary We are looking for a GTM DevOps Engineer to join our Business Systems team and own the reliability, automation, and delivery infrastructure behind our Go-To-Market (GTM) technology stack. This role sits at the intersection of platform reliability and CI/CD engineering, ensuring that our critical business systems — including Salesforce, NetSuite, MuleSoft, Workato, and an expanding portfolio of AI-powered workloads — are deployed consistently, operate resiliently, and scale with the business. You will partner closely with Business Systems developers, architects, and business stakeholders to build and maintain the pipelines, monitoring frameworks, and operational standards that keep our GTM systems healthy and our release cycles fast and predictable. As our team builds and deploys AI agents across GCP Cloud Run and AWS Bedrock AgentCore, you will serve as the infrastructure and deployment owner for these workloads — bringing engineering discipline to an environment where AI-generated code is increasingly entering production. This is a hands-on engineering role for someone who thrives in complexity, takes ownership of platform uptime, and brings a software engineering mindset to business application operations — directly supporting GTMSOE's broader mission of operational excellence across the GTM org. Key Responsibilities CI/CD & Release Engineering Design, build, and maintain CI/CD pipelines for Salesforce (SFDX/Salesforce CLI), NetSuite (SuiteScript/SuiteBundler), MuleSoft (Anypoint Platform), and Workato; establish branching strategies, environment promotion standards, and release gatin

PythonNode.jsAWSGCP
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 The Mission The Foundry is ClickUp's internal AI innovation lab — embedded inside GTM Systems and accountable for turning AI capabilities into production-grade, internally deployed products that make every GTM function faster and smarter. We build the infrastructure that powers AI-first work across Sales, Marketing, Post-Sales, and Revenue Operations. As the Senior Software Engineer on this team you will own the technical delivery of our MCP server platform, agent orchestration layer, and internal tooling — shipping production systems used daily by hundreds of ClickUp employees, and scaling your own throughput by treating AI tools as first-class engineering collaborators. What You'll Own MCP Server Platform Design, build, and operate Model Context Protocol servers that expose CRM, ticketing, analytics, and communication data to AI agents across the GTM stack Implement Okta PKCE authentication flows and RBAC policy enforcement so agents access only the data they're authorized to touch Maintain deployment infrastructure on AWS (Bedrock, Lambda, ECS, API Gateway) and contribute to GCP workloads where applicable Own observability: structured logging, distributed tracing, latency SLOs, and on-call runbooks for every production server Agent Orchestration & AI-Native Products Build and maintain multi-step autonomous agents that execute end-to-end GTM workflows — lead qualification, deal room assembly, onboarding automation, support triage, and more Architect prompt engineering frameworks, tool-call schemas, and agent evaluation harnesses that make AI behavior predictable and auditable Integrate with LLM p

TypeScriptPythonReactNode.js
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg

TypeScriptPythonGCPKubernetes
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Help power the development of Replit Agent as an engineer in the Replit Cloud organization. The Replit Cloud team builds Replit’s first party cloud infrastructure so users can build, scale, and succeed entirely on Replit. They manage databases, application storage, app publishing and hosting, development/production environment splitting, custom domains, and more. By having a set of first party services that integrate seamlessly, you will power one of Replit’s key product differentiators. You will: Work closely with designers and product managers, to quickly iterate on Replit Cloud to continually grow and improve the product. Drive full-stack feature development from conception to deployment, taking ownership of key product initiatives. Contribute to architectural decisions that shape the future of our product. Ship product and build infrastructure as a true full stack builder using: TypeScript, React, CSS, Postgres, Go, and Terraform. Examples of what you could do: Leverage our unique cloud infrastructure to build differentiated full product experiences, helping non-technical or semi-technical users remove roadblocks to success. Leverage AI agents to proactively optimize or suggest app improvements on latency, reliability, SEO, and more. Be part of engineering leadership, steering teams towards the highest impact work and supporting initiatives across the company. Required skills and experience: Bachelor’s degree in Computer Science or related field, OR equivalent real-world experience in engineering roles. Comfortable building with our tech stack: TypeScript, React, Go Preferred Qualifications Experience building user facing platform as a service products. Experience with AI/agentic systems. Previous e

TypeScriptReactAIGo
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you. In this role you will: Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently. Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base. Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services. Required skills and experience: Distributed systems: Track record of working with platform-as-a-service, distributed storage, o

GCPLinuxAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching

AWSAzureRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea

AWSAzureKubernetesCI/CD
O
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the role As a foundational FDE manager, you’ll lead FDE through high-stakes, ambiguous customer deployments and own technical and business value outcomes end to end. You’ll grow a team that can operate under pressure and help OpenAI learn from the field. You’ll partner closely with Product, Research, Sales, and GTM to ensure fieldwork informs roadmap priorities, drives new exploration, and supports safe deployment at scale. Your decisions will influence how OpenAI is trusted by the customers closest to our deployment work. Your success will be measured by how consistently your team ships, how clearly you deliver signal to Research and Product, and how durable your team and delivery model prove to be. This role is based in New York City We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. This role also will require travel up to 25%. In this role you will Lead and grow a team of FDE delivering production systems with frontier models Own end-to-end delivery outcomes through clarity, speed, tight coordination, and technical quality Codify what works into tools, playbooks, and roadmap inputs that create leverage for both OpenAI and our wider developer community Notice early indicators and raise them with urgency, whether in product behavior, customer environments, or delivery practices Use judgement to distinguish what requires action and what does not Set a high bar for FDE performance and support each person’s growth through direct, actionable feedback Define how we staff and support field teams that can scale without added complexity You might thrive in this role if you Bring 8+ years of engineering or technical delivery experience, including 2+ years managing high-performing FDE or custo

JavaScriptPythonJavaAWS
🔔

Get new ai deployment engineer jobs in United States by email

Daily job updates · Unsubscribe anytime