Become a part of our caring community You have shipped AI products before. You understand the difference between a demo and a production system. You have strong opinions about evaluation frameworks because you have experienced the consequences of operating without them. You are at your best when you own architecture decisions while continuing to build and deliver critical code yourself. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations back to source documents, and route complex cases to human experts. The output of these systems supports healthcare decisions that impact real members. As a Lead AI Applied Engineer, you will provide technical leadership for AI-enabled products and platforms, define architectural direction, establish engineering standards, and personally design and build the most critical components of our systems. You will lead through both technical expertise and execution, helping the team deliver reliable, scalable, and auditable AI solutions in a highly regulated healthcare environment. Why Join Us Lead the architecture of production AI systems where LLMs are foundational to the product experience. Make key technical decisions regarding model selection, system boundaries, platform architecture, and build-versus-buy strategies. Own the highest-risk and highest-impact technical challenges involving reliability, explainability, and correctness. Influence engineering culture and establish standards that shape how the team builds and ships AI products. Work on systems operating at meaningful scale, processing millions of documents and supporting healthcare decisions across a large member population. Partner
Jobs in United States
Production Director in United States
1,337 active opportunities · Updated October 2026
Showing
15 jobs
Explore current production director jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
As a Product Marketing Manager for NVIDIA NIM, you will play a pivotal role in helping enterprises deploy AI. NVIDIA NIM is revolutionizing AI from experimentation to production, packaging optimized, production-ready models for rapid deployment. Our product marketing team operates at the intersection of developer tooling and enterprise go-to-market, ensuring that our narrative resonates with ML engineers and CTOs alike. If you relish diving into technical depth and translating it into compelling stories, this is the perfect opportunity! What you will be doing: Launching products: You'll build and manage launch plans for high-visibility NIM releases, coordinating assets, timelines, keynote slides, demos, and press materials while driving cross-functional execution with product management, engineering, technical marketing, and PR or equivalent experience. You'll align the team on messaging and positioning for NIM across developer and enterprise audiences. You will develop messaging docs, sales enablement materials, customer presentations, solution overviews, and web content. Driving awareness: You'll identify target audiences and content gaps, then build the assets that fill them — blogs, webinars, demos, solution briefs, and more. Crafting the ecosystem story: You'll drive co-marketing engagements with NVIDIA's model providers, cloud partners, and ISV ecosystem to showcase the full range of possibilities with NIM. Collaborating with PR: You'll work with PR on press launches to ensure the NIM story is accurate, compelling, and consistent across every channel. What we need to see: Excellent written and verbal communication skills, with a proven track record of articulating technical value to both developer and executive audiences. College degree or equivalent experience. More t
As a Senior Backend Engineer on Coder's Enterprise Experience team, you'll build the systems that help large organizations run Coder in production with confidence. You'll improve how Coder scales, how it's upgraded, and how reliably it performs in regulated, air-gapped, and enterprise environments. You'll work on a genuinely cross-functional team of backend, platform, and QA engineers who own the end-to-end experience for Coder operators. From designing new features to evolving Coder's architecture, you'll partner across engineering and product to solve complex problems and ship software that operators trust. What you'll do here Design and build new features end to end, from technical design through production rollout. Design and implement backend architecture changes that support Coder's long-term scalability goals. Investigate and resolve scalability bottlenecks under production-like load, from database access patterns to concurrency handling in coderd. Improve database migration safety and upgrade reliability through schema compatibility, background migrations, and safe rollback strategies. Own the backend side of issues surfaced by Coder operators and administrators. Document the design, implementation, and operational tradeoffs of the systems you build. Participate in code reviews, RFC-style design discussions, and on-call rotations for the services you own. What we're looking for 5+ years of professional software engineering experience, including significant production experience with Go. Deep understanding of Go's concurrency model, including goroutines, channels, the sync package, and debugging race conditions under real-world load. Experience designing and operating relational databases in production, including schema design, migrations, and transactions. Strong verbal and written communication skills. Exceptional debugging and troubleshooting skills, with the persistence to drive complex problems to resolution. A self-motivated, analytical engineer who enj
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We believe Plaid has the power to be the next-gen Credit Bureau - supporting large scale adoption of cash flow into the credit underwriting process. The Credit Decisioning platform team is responsible for building best-in-class cashflow based insights products that enable lenders to make more holistic lending decisions and empower broader access to Credit products for prospective borrowers. We own the systems and tooling that form the platform to build and serve these insights at huge scale, partnering with our Data partners to release new products yearly. You will be defining the future architecture of Credit insights products and executing against an ambitious product roadmap. You will partner with our Product, Data Science, and Machine Learning team to iterate on and productionize new insights that enable our customers to make more holistic lending decisions. Responsibilities: Leading technical architecture and execution across credit insights products: everything from data fetching and online feature serving for API requests, to offline production pipelines and tooling for model training. Scaling and evolving the architecture through an expected ~100x increase in load from deterministic factors
About the team The OpenAI for Government team partners with federal, state, local, defense, national security, and international public-sector organizations to securely and responsibly adopt frontier AI, strengthen public services, and deliver meaningful mission impact. About the role OpenAI is seeking a strategic and deeply technical leader to serve as the Head of Government Technical Success. This leader will oversee the technical functions across the government customer lifecycle, spanning pre-sales engagement, prototype-to-production delivery, and post-sales adoption and value realization. You will define and operate a unified technical success strategy across federal civilian, defense and national security, state and local, international public sector, and industry partners. Your mission is to help customers identify the highest value applications of OpenAI’s technology, navigate the technical and organizational requirements of government environments, move those applications into production, and scale adoption in ways that deliver measurable mission impact. This role combines organizational leadership, technical judgment, executive customer engagement, operating rigor, and product influence. You will partner closely with Government Sales, Product, Engineering, Research, Security, Legal, Global Affairs and Policy, and other teams to ensure a seamless customer experience for governments in the United States and around the world. This role is based in Washington, DC. We offer relocation support to new employees. In this role, you will Set and continuously refine the strategy, operating model, and priorities for Government Technical Success, aligning the organization to OpenAI’s broader objectives and the distinct needs of government customers. Build and lead an organization of technical success personnel, including hiring, organizational design, manager development, career growth, and high standards for technical and customer-facing excellence. Create a seamless
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Infrastructure Security Engineer to design and secure the core systems that power our platform. This role focuses on building security directly into our infrastructure—from container isolation and orchestration to identity and secrets management in a multi-tenant, cloud-native environment. You’ll work closely with engineering teams to define secure primitives and ensure our platform is resilient, scalable, and trustworthy by design. This is a hands-on, deeply technical role focused on real systems, not compliance or policy. What You'll Do: Platform & Runtime Security Design and improve isolation mechanisms for multi-tenant workloads (containers, sandboxing, execution environments) Strengthen boundaries between customers, workloads, and internal systems Identify and mitigate risks in distributed, dynamic compute environments Container &
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role Is Different This is not a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers on problems that push LLMs to their limits. You’ll rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. Train and customize frontier models — not just use APIs. You’ll leverage Cohere’s full stack: CPT, post-training, retrieval + agent integrations, model evaluations, and SOTA modeling techniques. Influence the capabilities of Cohere’s foundation models. Techniques, datasets, evaluations, and insights you develop for customers will directly shape the next generation of Cohere’s frontier models. Operate with an early-startup level of ownership inside a frontier-model company. This role combines the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab. Wear multiple hats, set a high technical bar, and define what Applied ML at Cohere becomes. Few roles in the industry combine application, research, customer-facing engineeri
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role Is Different This is not a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers on problems that push LLMs to their limits. You’ll rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. Train and customize frontier models — not just use APIs. You’ll leverage Cohere’s full stack: CPT, post-training, retrieval + agent integrations, model evaluations, and SOTA modeling techniques. Influence the capabilities of Cohere’s foundation models. Techniques, datasets, evaluations, and insights you develop for customers will directly shape the next generation of Cohere’s frontier models. Operate with an early-startup level of ownership inside a frontier-model company. This role combines the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab. Wear multiple hats, set a high technical bar, and define what Applied ML at Cohere becomes. Few roles in the industry combine application, research, customer-facing engineeri
From $183K/yr
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. The Opportunity The Autonomous Freight Systems team is a brand new, AI-first engineering team in San Francisco, building Flexport's client-facing rates platform and self-serve freight booking experience from the ground up. We own two of the highest-leverage surfaces in the Client App: how clients see pricing and how they book freight without manual intervention. As a Senior Engineer, you will own significant technical domains end-to-end, not just features, but whole problem spaces. You will drive the engineering of systems that power rate visibility, pricing intelligence, and AI-assisted booking across ocean, air, and trucking. You don't wait to be told what to build next; you identify the highest-leverage problems, propose solutions, and see them through to production. You will be a technical force multiplier on the team: raising the bar on code quality, helping junior engineers grow, and shipping with a velocity and confidence that comes from deep ownership. This is a ground-floor opportunity to build the platform that moves Flexport from an assisted-sales model to a tech-run one for the long tail of our client base—turning the complexity of global trade into a s
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver effective performance ads to our users, and more business values to our advertisers. We’re looking for an EM to lead a team of exceptional ML infrastructure engineers, build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You Will: Lead strategic planning and roadmap execution of scalable production-ready ML systems including model training, data pipelines, feature engineering and model inference. Own the architecture, establish engineering best practices of scalability, reliability, and cost-effectiveness of ML infrastructure (e.g., training, serving, feature). Work closely with data scientists, ML engineers, platform teams, and product stakeholders to design, implement, and operate robust ML platforms that accelerate model development and deployment. Recruit, mentor, and grow a high-performing team of ML infrastructure engineers. You Have: 5+ years of experienc
From $326.1K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads & Discovery business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver more value to our users and our advertisers. As a Machine Learning Infrastructure Engineer, you’ll build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You will: You will co-design models and systems, working at the intersection of model architecture and ML infrastructure, partnering closely with core modelers, data and AI infrastructure engineers, and product teams to push the boundaries of large-scale training and serving. Your work will span recommendation, search, and agentic applications, including large transformer architectures, LLMs, generative rankers, and efficient offline and online content-understanding systems. You will investigate model, data, and systems tradeoffs end to end—from data pipelines and distributed training to low-latency inference and production serving. This includes designing efficient KV-cache strategies, applying p
From $192K/yr
Applied AI is where Datadog's ambitious AI bets get built and shipped ( Bits Chat , updog ). We sit at the intersection of research and product: turning promising capabilities from Datadog AI Research lab and the research community into production systems that reach real customers. The team builds the foundations for agentic systems capable of operating at scale in complex production environments. Current bets span agents that run autonomously at scale, context and memory layers that make those agents more intelligent over time, and tools that help customers build and validate AI-native services in production. The mandate is to move fast from idea to customer impact, and when a product finds its footing, to set it up for growth. As an Engineering Manager I in Applied AI, you will lead a team of engineers and applied scientists working on one of these challenges. You will define technical direction, run short feedback loops, make deliberate decisions about what to pursue or stop, and work closely with product managers, research teams, and cross-functional partners to ship AI capabilities that matter. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Lead and develop a team of engineers and applied scientists focused on building the foundations for agents operating at scale Work closely with product managers, research teams, and cross-functional partners to shape the team's bets from initial framing through to broader adoption, with a clear definition of success criteria at each stage Own end-to-end delivery of high-quality AI systems, from early research exploration to production-grade reliability, with high standards for operational excellence, system reliability, and technical quality Navigate the unique challenges of shipping AI-powered products: balancing quali
From $234K/yr
The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs) and generative AI. We provide robust, scalable observability for AI workloads, including drift detection and model evaluation, and behavior tracing, enabling customers to ship AI with confidence. As a Staff Engineer, you’ll lead the development of new features and foundational capabilities within Datadog’s LLM Observability product. You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems in the fast-moving AI landscape. Your work will directly impact how our customers monitor, troubleshoot, and optimize LLM-based applications in production. Join us in building the foundational tools that make AI systems observable, understandable, and reliable in the real world. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive design and implementation of LLM observability features. Ideate, prototype, and scale new product features to provide insights and drive improvements for generative AI systems Work cross-functionally with other eng teams, product, UX, and applied science to iterate fast and find product-market fit Develop and extend tools for tracing, evaluating, and debugging LLMs Influence architecture decisions and mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities Stay current with industry trends and advancements in machine learning and observability, driving innovation within the team Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or r
Other cities to consider
More places hiring for this role
Get new production director jobs in United States by email
Daily job updates · Unsubscribe anytime