Jobiba hiring network

Ai Deployment Manager Jobs

10,000 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current ai deployment manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. To be clear, this is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects. Drive customer impact by designing, implementin

pythondockermachine learning
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Infrastructure Software Engineer at Baseten, you'll build and maintain components of our ML inference platform that powers production AI applications. You'll contribute to the core infrastructure, enabling developers to deploy, scale, and monitor ML models with high performance. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Develop infrastructure components for our ML inference platform using Python and Go Implement and maintain Kubernetes deployments for model serving Contribute to our inference orchestration layer for model deployments Build and enhance monitoring systems for model performance metrics Implement efficient resource management solutions for ML workloads Support infrastructure automation to improve ML deployment workflows Work closely with team members to implement technical solutions Help balance performance optimization with system reliability Participate in technical discussions around infrastructure improvements Learn and apply infrastructure best practices REQUIREMENTS Bachelor's degree or higher in Computer Science or related field Proficient coding abilities in one or more popular programming or scripting languages; Go proficiency is a plus Working knowledge of Kubernetes and containeriza

pythonkubernetesrest
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le

kubernetesci/cdrest
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr

kubernetesrestmachine learning
View job →
R
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg

typescriptpythongcp
View job →
R
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg

typescriptpythongcp
View job →
R
Replit
📍 Foster City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you. In this role you will: Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently. Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base. Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services. Required skills and experience: Distributed systems: Track record of working with platform-as-a-service, distributed storage, o

gcplinuxai
View job →
N
Nuro
📍 Mountain View• Full-time• From $183.8K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team Our robotics team is growing and we are looking for an ML Software Engineer to join our Online Mapping team. We are searching for an engineer with robotics and machine learning expertise to work on challenging problems in the design and implementation of onboard mapping models and algorithms, data and label management, as well as training pipelines. About the Role Building robust ML and/or mapping systems that work with real data (cameras, LiDAR, etc.) in uncertain environments, and a strong desire to contribute to the future of robot navigation for logistics and transportation. About the Work Research, develop, and implement state-of-the-art online mapping models and algorithms. Analyze and characterize the performance of the online mapping system, identifying opportunities for architecture, data or evaluation improvements in a E2E ML system. Work cross functionally with other ML teams to integrate our models

pythonmachine learningai
View job →
N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role We are looking for talented engineers to join our Performance team to optimize the performance of Nuro’s AV software, ensuring our vehicles can react quickly and safely to the world around them. The team builds systems and tools for continuous performance analysis and drives latency reduction and resource efficiency efforts to ensure the autonomy teams can implement an autonomy stack that is efficient and performant for the current and future generations of the Nuro Driver. About the Work Analyze, profile, debug, monitor, and optimize the performance of AV software Design and develop systems and tools for memory management, thread prioritization, process/thread lifetime management Work with engineers from different teams to define the system-level architecture and building blocks Build core libraries and APIs to enable autonomy engineers to write high-performance code Drive and encourage best practices within the team and t

reactrestai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role We are looking for talented engineers to join our Performance team to optimize the performance of Nuro’s AV software, ensuring our vehicles can react quickly and safely to the world around them. The team builds systems and tools for continuous performance analysis and drives latency reduction and resource efficiency efforts to ensure the autonomy teams can implement an autonomy stack that is efficient and performant for the current and future generations of the Nuro Driver. About the Work Analyze, profile, debug, monitor, and optimize the performance of AV software Design and develop systems and tools for memory management, thread prioritization, process/thread lifetime management Work with engineers from different teams to define the system-level architecture and building blocks Build core libraries and APIs to enable autonomy engineers to write high-performance code Drive and encourage best practices within the team and t

reactrestai
View job →
M
Mongodb
📍 Sydney• Full-time
1mo ago

We’re looking for a Software Engineer 3 to help bring Voyage’s embedding models - used for semantic search, retrieval, and AI-native experiences; to the platforms and environments where customers already run their workloads, beyond first-party MongoDB Atlas. You’ll join the broader Search and AI Platform organization and collaborate closely with the engineers building Voyage’s first-party inference. Together, we’re extending that platform across cloud marketplaces, third-party inference providers, and self-managed deployments so customers get the same Voyage models, behaving consistently, wherever they choose to run them. As a Software Engineer 3, you'll focus on building the systems, tooling, and deployment workflows that power third-party model delivery. You'll own key components of how Voyage models are packaged, validated, and deployed, work across teams to ensure tight integration with the core inference platform, and contribute to delivery surfaces designed for reliability, observability, and ease of use. We are looking to speak to candidates who are based in Sydney for our hybrid working model. What you'll do Port and tune the model server that runs Voyage embedding and reranking models: improving inference performance, consistency, and runtime behavior across environments Productionize new Voyage models for delivery beyond first-party Atlas, owning the packaging, configuration, and deployment workflows that get them running on AWS, Azure, GCP and more Design correctness, correlation, and performance validation that proves third-party deployments match first-party behavior Build operability into every surface: structured logging, metrics, diagnostics, and health checks with tools like Prometheus and OpenTelemetry Debug problems that span model servers, containers, deployment configuration, and partner cloud environments Work alongside Voyage's model-serving teams, and partner with GTM, SAs, TSEs, and strategic customers on the hardest external deployments Who

pythonmongodbaws
View job →
O
1mo ago

About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) which benefits all of humanity. This long-term undertaking brings the world’s best scientists, engineers, and business professionals into one lab together to accomplish this. In pursuit of this mission, our Go To Market (GTM) team is responsible for helping customers learn how to leverage and deploy our highly capable AI products across their organization. The team comprises Sales, Solutions, Support, Marketing, and Partnership professionals who collaborate to create valuable solutions that will help bring AI to as many users as possible. About the Role Our Government Sales team has a unique mission to help the U.S. Intelligence Community understand the transformative impact that highly capable AI models can bring to its most critical missions. This role combines technical understanding, strategic vision, relationship management, and value-driven sales strategy tailored specifically to Intelligence Community customers. You’ll drive key opportunities throughout the full sales cycle, from pipeline generation through deployment and expansion. You’ll collaborate closely with researchers, engineers, solution strategists, and cross-functional partners to help Intelligence Community customers advance their missions through AI. This role is based in Washington, D.C. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In addition, this role requires frequent on-site engagement at customer and partner facilities across the Washington, D.C./Northern Virginia corridor, including classified environments. In this role, you will: Manage a focused set of key U.S. Intelligence Community accounts, developing and executing comprehensive account strategies. Lead Intelligence Community customers through their AI adoption journey, from initial consideration through successful deployment and expansion. Build trusted relationships with senior

awsrestai
View job →

About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a new team within USRO focused on building operational capacity for new, ambiguous, or fast-moving areas of work. The team helps define what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. As a Strategic Operations Lead, you will focus on large, cross-functional initiatives that require clear thinking, technical fluency, strong execution, and the ability to bring structure to undefined problems. About the Role We are seeking a Strategic Operations Lead to drive new and existing strategic operating builds across User Safety & Risk Operations. This is a senior IC role for someone who can turn broad, undefined priorities into clear operating models, launch plans, requirements, stakeholder alignment, documentation, reporting, and execution rhythms. This role will often support initiatives where OpenAI is developing new products or partnerships and the operating model is still being defined. These programs have a direct user safety and risk nexus because new deployment models can change what signals OpenAI can see, who owns response decisions, and how user-impacting risks are detected, escalated, and resolved. You will clarify what OpenAI owns, what partner teams own, what signals we can reliably monitor, how issues should be escalated, and how the workflow should evolve from launch support into a durable operating model. The right person is highly strategic and deeply practical. They can move from executive-level framing to detailed workflow design, stakeholder management, SOPs, launch readiness, ri

sqlawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea

awsazurekubernetes
View job →
O
OpenAI
📍 Singapore• Full-time
1mo ago

About the team OpenAI’s Enterprise Go-To-Market organization helps the world’s largest companies adopt and scale our AI platform across their business—from ChatGPT Enterprise to our developer platform and APIs. We partner with organizations across Financial Services, Life Sciences, Retail, Manufacturing, Technology, and other key industries across Asia Pacific to build AI-powered customer experiences, transform operations, and accelerate innovation. As adoption accelerates through pilots, experimentation, and developer-led use cases, the GTM team turns that momentum into durable, enterprise-wide deployments. Sales Development sits at the front of this motion—where technical curiosity becomes executive engagement and OpenAI’s enterprise relationships begin. About the role We are hiring a APAC Sales Development Leader to build and lead OpenAI’s enterprise Sales Development organization across Asia Pacific. Reporting directly to the Head of Global Sales Development, this role is a core part of the Enterprise GTM leadership team. This is a high-impact leadership role focused on scaling pipeline generation and market development across a diverse and rapidly growing region. You will define how OpenAI engages enterprise customers across APAC, build regional SDR leadership and teams, and create the operating model that converts product interest, developer activity, and inbound demand into high-quality pipeline for our Enterprise and Strategic Accounts organizations. In this role, you will Build and scale OpenAI’s APAC Sales Development organization, including hiring strategy, org design, enablement, and performance management. Lead and develop front-line SDR managers and SDR teams across key APAC markets. Own the regional enterprise pipeline generation strategy, ensuring consistent, high-quality opportunity creation across the region. Translate product usage, pilots, developer engagement, and inbound interest into qualified enterprise opportunities. Develop localized and ve

awsrestai
View job →
🔔

Get new ai deployment manager jobs by email

Daily job updates · Unsubscribe anytime