ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Forward Deployed Engineers work directly with the largest and fastest-growing AI companies in the world, owning their technical outcomes on Baseten and taking on the hardest problems in serving and improving models at scale. The work spans the model lifecycle: inference, post-training, and the systems that tighten the loop between them. Act as each account's de facto CTO on Baseten, with final accountability for how their workloads are designed, run, and scaled. Take customer objectives from vague to shipped: frame the problem, define the spec and success criteria, build the PoC, and carry it through to production quickly, using the right tools for the problem. Design the evals and benchmarks that isolate where quality or performance falls short, then close the gap yourself, whether that means optimizing inference, improving the model through post-training, or reworking the eval itself. Be the first responder to mission-critical failures including triage, owning the fix directly or route to the owning team and stay accountable until it ships. Build internal systems so that each engagement is faster than the last. This includes tooling and automation for eval and deployment infrastructure, and the recipes and reference implementations that make the product more self-serve. Shape the product itself, channeling what your accounts need into the roadmap and shipping fixes and features into Baseten's codebase yourse
Jobs in United States
Ai Deployment Manager in United States
5,082 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai deployment manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As an early member of Baseten's Platform Team, you will be pivotal in building internal infrastructure to support our engineering organization. You will own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across Baseten to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. If you are passionate about elegant solutions—like streamlined monorepos, lightning-fast CI pipelines, and thoughtfully designed shared libraries—you'll thrive at Baseten. RESPONSIBILITIES Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces,
We are looking for a Senior Forward Deployed Engineer to join the Customer Solutions team. You will be the technical authority embedded with our most complex customers, guiding them through deployment, architecture, onboarding, and the adoption of agentic development workflows. You bring deep hands-on experience from prior roles and use that depth to advise, design repeatable patterns, and drive customer outcomes end to end. You operate autonomously, own the technical success of your customers, and bring their experience back to shape how Coder builds and delivers. This is not an execution-only role. You are equally comfortable doing deep technical work with a customer and stepping back to design the repeatable pattern behind it. You are energized by ambiguity, motivated by customer outcomes, and capable of influencing organizational change alongside the technical work. This position is required to sit in the Eastern Time Zone. What You'll Do Serve as the primary technical authority for post-sales customers, guiding deployment architecture, environment design, and adoption of Coder across both human and AI development workflows Own onboarding engagements end to end, ensuring customers move from contract to productive adoption with speed and confidence Lead Get Well engagements where architecture decisions, rollout patterns, or organizational dynamics are limiting customer health or growth Help customers implement the technical and organizational changes required to adopt agentic development practices at scale Design and document repeatable delivery patterns across onboarding, architecture, and adoption that can scale across customer segments Design and recommend reference architectures tailored to each customer's cloud environment, security posture, and organizational constraints Translate customer environment complexity into clear guidance on networking, ingress, identity, and infrastructure patterns Anticipate technical and operational risks, escalate to the right
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role At WRITER, we're not just building the future of enterprise AI; we're actively orchestrating it. The director, solutions architecture plays a pivotal role in this journey, leading a team of exceptional solutions architects who empower our enterprise customers to unlock the full potential of our cutting-edge generative AI platform. You’ll be at the forefront of driving AI transformation, moving beyond theoretical discussions to hands-on deployment and measurable impact. This is your chance to lead, mentor, and inspire a team dedicated to solving complex business challenges with innovative AI solutions, directly influencing customer success and WRITER's rapid growth. This is a full-time, hybrid role based out of our San Francisco hub. You will report directly to the VP of solutions architecture. 🦸🏻♀️ What you'll do Lead and empower a team of highly skilled solutions architects, fostering their technical growth and career development across complex enterprise AI engagem
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi
About the Team The Enablement Lead (EL) team enables organizations to turn OpenAI products into real, sustained impact through world-class enablement and training execution. Our mission is to help customers successfully adopt and operationalize AI across their organizations. We partner with enterprises to translate the potential of OpenAI’s technology into durable capability—through structured training, technical enablement, and scalable deployment programs. By helping customers move from experimentation to production, the EL team accelerates time-to-value, deepens product adoption, and helps make OpenAI indispensable to how organizations work. About the Role The Enablement Lead, Builder role is a specialist post-sales technical enablement role focused on delivering high-impact enablement and adoption services across OpenAI’s product suite. You will design and deliver technical learning experiences covering OpenAI APIs, Codex, agents, evaluations, and related platform capabilities. You will work with engineers, AI and platform teams, administrators, security stakeholders, product leaders, and executive sponsors. This role blends deep technical fluency, instructional design, and customer advisory. You will lead live trainings, workshops, and adoption interventions for audiences ranging from hands-on builders to executive leaders, helping customers understand not just what OpenAI’s products can do, but how to use them effectively in real-world contexts. Success in this role means accelerating customer confidence, increasing product adoption, helping customers progress toward production use, and turning lessons from individual engagements into resources and practices that benefit many customers. This role is based in our San Francisco HQ. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the technical enablement of OpenAI products, including OpenAI APIs, Codex, Agents, Evaluations a
About the Team OpenAI’s Business organization works with customers and partners on some of our most complex and consequential opportunities. These efforts require rigorous strategy, strong operating leadership, and coordinated execution across commercial, product, technical, deployment, and go-to-market teams. About the Role We are hiring a Business Lead to serve as the operating leader for a strategically important initiative anchored in a major partnership. This person will turn an ambitious, cross-functional mandate into a clear strategy, operating plan, decision structure, and set of measurable outcomes. This role combines strategy and operations, product judgment, commercial skills, and the ability to get things done. You will identify the most promising product and customer opportunities, develop a point of view on how our products should work together, shape the commercial approach, and personally drive execution across both organizations. This is a hands-on role for someone who wants to own outcomes, not just coordinate work. You will structure ambiguous problems, establish priorities, build trusted relationships with internal and external stakeholders, and work across product, engineering, deployment, and go-to-market teams to remove blockers and deliver results. Success means establishing a durable operating model for the initiative, improving decision velocity, translating strategy into coordinated execution, launching a joint go-to-market motion, landing an initial cohort of customers, and delivering a high-quality enterprise deployment. In this role, you will: Own the initiative’s integrated strategy and operating plan, including priorities, desired outcomes, metrics, owners, dependencies, and decision points Structure complex and ambiguous business problems, develop fact-based recommendations, and translate them into clear choices and executable plans Define success measures and build operating reviews that surface progress, risks, tradeoffs, and requi
About the Team OpenAI is evaluating multiple infrastructure pathways, including powered land, colo/BTS, and NeoCloud opportunities. The Site Readiness & Development team provides the diligence layer needed to compare opportunities, identify risk, and support credible deployment decisions across those pathways. About the Role The NeoCloud & Colo Due Diligence Lead will evaluate third-party infrastructure opportunities where OpenAI is considering deployment through NeoCloud, colo, or BTS structures. This role will focus on facility and deployment readiness, including MEP readiness, rack strategy, developer capability, facility design, power deliverability, schedule credibility, and operating assumptions. Unlike the land diligence team, this role is centered on technical and operational readiness of third-party infrastructure rather than greenfield site master planning, civil development, and entitlement strategy. This is an individual contributor lead role and does not have direct reports initially. The role determines whether each opportunity is fit-for-use and fit-for-service against OpenAI facility, rack, power, network, reliability, and operational standards; identifies material deficiencies and tracks remediation with developers/operators; and evaluates commissioning, validation, AHJ/code, and deployment interfaces such as structured cabling, network readiness, and high-density rack support where relevant. Key Responsibilities Lead diligence on NeoCloud, colo, and BTS opportunities across technical and operational readiness dimensions. Assess each opportunity against OpenAI facility, rack, power, network, reliability, and operational standards to determine deployment fit. Validate MEP readiness, rack deployment strategy, facility design assumptions, power deliverability, and schedule credibility. Identify material deficiencies and work with developers/operators to define remediation plans, owners, timing, and residual risk. Review reliability, availabilit
About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of AI at unprecedented scale. Through a combination of strategic partnerships and self-built campuses, we are developing and operating large-scale data center infrastructure across power, cooling, networking, compute, construction, and site operations. The scale and complexity of this infrastructure introduces a broad range of environmental, health, and safety considerations across site development, design, construction, equipment deployment, commissioning, and ongoing operations. EHS is a critical part of how we build infrastructure that is safe, resilient, compliant, and capable of operating at scale. About the Role We are seeking an EHS Lead to establish and drive environmental, health, and safety strategy across OpenAI’s rapidly expanding compute infrastructure portfolio. This role will develop the EHS framework for large-scale data center development and operations, partnering closely with engineering, construction, infrastructure delivery, facilities, operations, security, legal, environmental, and external development partners. The EHS Lead will help ensure that safety and environmental considerations are embedded into projects from early design and site development through construction, commissioning, and operations. The role will establish standards and operating mechanisms, assess and mitigate risks, oversee EHS performance across internal teams and third-party partners, and provide technical leadership on complex or high-consequence safety issues. Success requires the ability to operate strategically while maintaining strong technical depth and executional rigor in fast-moving, highly complex infrastructure environments. In this role, you will: Develop and own EHS strategy, standards, programs, and operating mechanisms across OpenAI’s data center and compute infrastructure portfolio. Establish scalable EHS requirements for site de
About the Team The Developer Experience team at OpenAI has a singular focus: empowering developers globally. Our mission is to provide every developer and startup on the planet with the most delightful and seamless experience to integrate AI into their applications and products. We ensure developers have the tools, resources, and support they need to unlock AI’s full potential. We create inspiring demos, developer tools, sample applications, and technical content that show developers how to build with Codex and frontier models like GPT-5.6, GPT-Live, and GPT-Image-2 to create powerful agents and AI-native applications. We collaborate closely with product, engineering, research, and GTM teams to ensure the developer journey, from onboarding with Codex to first API call to production deployment, is seamless, effective, and delightful. About the Role As a Developer Experience Engineer, you will create compelling technical content, developer tools, and sample applications designed to inspire developers and enable them to succeed with Codex and OpenAI’s APIs and products for developers. You will engage with developers and technical founders, demonstrating best practices and building innovative applications powered by frontier models, multimodal capabilities, and tools like Codex. We’re looking for people who combine strong technical skills, creativity, and a passion for engaging with and empowering developers. In this role, you will: Develop demos and sample applications that showcase best practices for building with Codex, frontier models, multimodal capabilities, and agents. Create high-quality technical content—including tutorials, blog posts, videos, and code samples—to educate and inspire the developer community about our models, APIs, and Codex. Actively engage with and foster a vibrant local and global developer ecosystem around OpenAI’s platform and products. Represent OpenAI at developer events and online, serving as a knowledgeable and approachable advocate for
About the Team The Consumer Products team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. Within Consumer Products, the camera stack is a critical sensing component. The team partners closely with electrical engineering, silicon vendors, systems, and higher-level perception and product teams to bring up new hardware, stabilize capture pipelines, and ensure camera systems are robust, debuggable, and ready for real-world deployment. This work spans early prototypes through production, with a strong emphasis on correctness, repeatability, and long-term reliability. About the Role As a Camera Firmware Engineer, you will own low-level camera enablement on custom hardware—from early board bring-up through stable production capture. You will develop and maintain the firmware and software that makes camera sensors reliable, controllable, and debuggable, forming the foundation for higher-level camera pipelines and product features. This role is highly hands-on and systems-oriented. You will work close to the hardware, diagnose real-world timing and integration issues, and build tooling that accelerates iteration across the entire camera stack. This role is based in San Francisco, CA. We follow a hybrid work model with four days per week in the office and offer relocation assistance to new employees. In This Role, You Will Bring up new camera sensors and modules on prototype and production boards, including link stability, sensor control, and correct power, reset, and clock sequencing. Develop and maintain low-level camera software, including sensor drivers, board configuration, and camera subsystem integration across hardware revisions. Enable and validate core capture paths for development and production, including RAW capture for debugging, still capture, and hardware-
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
About the Team OpenAI’s Infrastructure organization builds and evaluates the systems that power advanced AI workloads. We work closely with hardware, modeling, and architecture teams to ensure that new platforms deliver real-world performance aligned with workload needs. Our team focuses on understanding workload behavior across evolving hardware platforms—bridging the gap between theoretical capability and observed system performance. About the Role We are seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by porting benchmarks and real-world workloads, analyzing performance, and identifying system bottlenecks. In this role, you will bring up workloads on new systems, characterize performance behavior, and adapt workloads to better utilize hardware capabilities. You will play a critical role in validating new platforms and ensuring that performance aligns with expectations across compute, memory, and networking subsystems. This role requires strong hands-on experience with performance analysis, workload optimization, and system-level debugging across hardware and software boundaries. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Port and enable benchmarks and real-world workloads on new hardware platforms. Evaluate system performance across compute, memory, storage, and networking subsystems. Identify and analyze performance bottlenecks and inefficiencies. Adapt and optimize workloads to better utilize hardware capabilities. Develop and run performance experiments and profiling workflows. Compare expected vs. observed performance and provide feedback to: hardware architecture teams performance modeling teams system and software engineers. Debug issues across the stack, including software, runtime, and hardware interactions. Provide actionable insights to guide platform readiness and deployment decisions. Qualifications E
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Vice President, Boise Front End Manufacturing Operations Location: Boise, Idaho Reports To: Senior Vice President, Global Front End Operations Position Summary Reporting to the Senior Vice President of Global Front End Operations, the Vice President, Boise Front End Manufacturing Operations is responsible for the overall leadership, performance, and strategic direction of Micron's Boise front-end manufacturing operations. This leader owns the safe, efficient, and predictable execution of all manufacturing activities while ensuring alignment with company objectives, technology roadmaps, customer requirements, and long-term capacity strategies. The Vice President serves as the senior operational leader for Boise manufacturing and is accountable for operational excellence across safety, quality, yield, cycle time, cost, delivery, workforce capability, and capacity utilization. As a member of the Global Front End Operations Leadership Team, the VP partners closely with Technology Development, Engineering, Business Units, Supply Chain, Procurement, Facilities, Quality, Finance, Human Resources, and peer manufacturing sites to drive world-class manufacturing performance and enterprise-wide outcomes. This leader will play a critical role in shaping the future of advanced semiconductor manufacturing through the deployment of next-generation technologies, automation, digital manufacturing capabilities, AI-enabled operations, and large-scale capacity expansions. Key Responsibilities<
Other cities to consider
More places hiring for this role
Get new ai deployment manager jobs in United States by email
Daily job updates · Unsubscribe anytime