ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues. Apply and scale optimization techniques across a wide range of ML models, particularly large language models. Collaborate with a diverse team to design and implement innovative solutions. Own projects from idea to production. REQUIREMENTS Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field. Experience with one
Jobs in United States
Model Designer in United States
2,223 active opportunities · Updated October 2026
Showing
15 jobs
Explore current model designer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto
About the Team The Human Data team at OpenAI is responsible for identifying and mitigating risks in advanced AI systems by designing evaluations, surfacing vulnerabilities, and collaborating closely with researchers to strengthen model reliability and public trust. About the Role As a Research Program Manager, you will lead initiatives that test the safety and robustness of OpenAI’s models through creative experimentation and structured evaluation. You’ll coordinate efforts across research and engineering teams to transform ambiguous risks into concrete research programs and influence future model development and deployment. We’re looking for people who are technically savvy, comfortable with ambiguity, and excited about shaping the future of safe AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead programs that explore unexpected model behaviors and identify failure modes. Translate vague or emergent risk signals into clear priorities and actionable research plans. Design and run creative evaluations, experiments, and red-teaming campaigns. Collaborate with research, product, and deployment teams to integrate findings into model training and deployment cycles. Develop repeatable systems for tracking model performance and understanding emerging behavior patterns. You might thrive in this role if you: Have strong experience in technical program management, with excellent organizational and communication skills. Are familiar with large language models, prompt engineering, or model evaluation techniques. Are comfortable managing fast-paced, high-uncertainty projects and shaping them from the ground up. Are creative and resourceful in devising new methods for testing model behavior and performance. Can effectively coordinate across technical and non-technical stakeholders to drive alignment and execution. About OpenAI OpenAI is an AI resear
About the Team The ChatGPT Model Flywheel team unified goal is to transform model advancements into great ChatGPT user experiences through reliable serving, rapid experimentation, safe deployment, and continuous improvement. Team Focus Areas Model Experimentation: Enable rapid, safe model validation for ChatGPT and Codex products through experiment automation and lifecycle management. Model Deployment: Ensure safe, scalable deployment of model capabilities with robust rollout and operational tooling. Automate capacity management and incorporate platform-wide health monitors. Model Measurement: Build comprehensive evaluation and measurement systems for model quality, from user signals to launch scorecards. Improve end-to-end feedback loops for continual model improvement. Key Partnerships Collaborate cross-functionally with teams including Model Measurement DS, Research, Codex, Fleet, Inference, and API. In this role, you will: Elevate and consolidate ChatGPT’s harness, context management, and system prompt frameworks. Drive expansion and improvement of multi-tier model experiences. Support and scale self-serve experiment capabilities and automated guardrails. Lead model rollout automation, capacity management, and health monitoring. Shape end-to-end measurement systems (evals, grader signals, user feedback, etc.). You might thrive in this role if you have: Proven experience leading engineering teams in complex, cross-functional environments. Demonstrated success shipping production systems at scale (ideally for AI or large backend services). Deep understanding of model-driven product development, deployment lifecycle, and measurement tooling. Excellent communication and collaboration skills—experience interfacing directly with engineering, research, and product stakeholders. Prior involvement with large language models, distributed infrastructure, or experimentation platforms is a plus. Why Work With Us Tackle highly impactful technical challenges at the cutting edg
Become a part of our caring community The Associate Vice President, Model & AI Governance, is the enterprise leader responsible for establishing and overseeing the organization’s framework for model governance, artificial intelligence (AI) risk management, responsible AI and AI governance. Reporting to the Chief Audit & Risk Officer, this executive provides independent second-line oversight and challenge of the organization’s use of models, advanced analytics, machine learning, generative AI, and emerging AI technologies. The role establishes the governance, risk-management, control, monitoring and escalation framework necessary to ensure AI models are deployed in a manner that is safe, ethical, transparent, explainable, compliant, secure and aligned with the organization’s mission and risk appetite. The Associate Vice President, Model & AI Governance, partners closely with executive leadership, technology, data and analytics, clinical/business leaders, compliance, legal, privacy, cybersecurity, information security, internal audit, enterprise risk management and other control functions to ensure AI-related risks are identified, assessed, governed, monitored, and appropriately reported. This leader will serve as an advisor to executive management on emerging model and AI risks, while maintaining appropriate independence from the teams developing and deploying models and AI solutions. Key Responsibilities Own and continuously enhance the enterprise model governance framework, including model identification, inventory, classification, risk tiering, development, validation, approval, implementation, monitoring, change management, retirement and documentation. Define model risk appetite, risk taxonomy, minimum control standards, governance requirements and escalation thresholds. Provide effective challenge over model
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion has been at the cutting edge of AI since before ChatGPT launched. The job of the Model Capabilities team is to keep us there. We own the model layer of Notion AI: integrating frontier models as they ship, keeping inference reliable and economical at scale, and building new capabilities that other teams take advantage of. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Bring new frontier models into production quickly, making them available for our users and our engineers. Make inference reliable: better error categorization, self-healing retries, and cross-provider failover. Own observability for the model layer, driving down both time to detection and time to fix. Build new model-level capabilities and help product teams adopt them. Act as connective tissue across Notion's AI teams: find the gaps, unblock people, and make sure fixes land with the r
About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT infrastructure team is responsible for ensuring that our products can serve rapidly growing demand with the performance, reliability, and quality our users expect. This work sits at the intersection of product demand, model deployment, inference, research, fleet, and capacity. The team translates changing product and model needs into clear capacity decisions and safe, scalable launches. About the Role We are seeking a Technical Program Manager to lead the operating system for Chat capacity and model deployment. You will connect demand forecasting and capacity allocation with model readiness, rollout planning, launch coordination, and post-deployment learning. You will also own mode deployment beyond capacity by working with cross functional teams across research, post-training, inference and product to own mainline model deployment. You will bring structure to constrained-capacity decisions, improve the tooling and mechanisms teams use to prioritize demand, and help new models reach users safely and efficiently. Success requires technical depth, sound judgment under ambiguity, and crisp execution across product, research, infrastructure, and operations teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own cross-functional programs for Chat capacity forecasting, allocation, headroom planning, and constrained-capacity operations. Build durable intake, prioritization, and decision mechanisms that connect product demand and model requirements to available serving capacity. Partner
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. Join NVIDIA's NIM team and be part of an exceptionally ambitious project in Santa Clara, CA! As a Senior Software Engineer, NIM Tools, you will have the remarkable opportunity to build a groundbreaking model customization and deployment lifecycle platform from inception. This isn't just another feature team—you will be defining the structure for a new product surface accessed by ISVs and CSPs internationally. Your work will empower customers to take models from selection through fine-tuning, evaluation, deployment, and compliance flawlessly. What you'll be doing: Compose and build the fine-tuning handoff pipeline, including LoRA adapter repackaging, re-quantization, and re-validation into NIM. Develop the evaluation harness, ensuring models meet our high standards. Implement the observability and attestation layer to produce auditable compliance artifacts. Work in close partnership with ISVs and CSPs to roll out NVIDIA NIMs on a large scale. Define and improve durable platform APIs, steering clear of one-off integrations. Ensure flawless completion of projects through strict attention to detail and proven methodologies. Wha
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new
About the Team The Integrated Marketing team sits at the center of product, brand, creative, media, research, and GTM work. We partner closely with PMM, Product, Developer Relations, Design, Comms, and agency teams to shape launches and campaigns that are clear, distinctive, and grounded in real audience insight. This is a growing function, so the team needs people who can both raise the quality of the work and build the operating model around it: bringing strong judgment, high agency, and trusted partnership to complex, visible moments. About the Role OpenAI business launches move quickly and often involve new products, model capabilities, or research that can change how businesses operate. We’re looking for an integrated marketing manager to help shape those launches from early strategy through execution. You’ll develop a point of view on the audience, positioning, creative direction, and channels, and help turn complex product and research advances into campaigns that connect with customers. You’ll partner across Product, Product Marketing, Research, Brand, Creative, Communications, Growth, and Sales, along with external agencies. The right person understands technology, has strong strategic and creative instincts, and knows how to bring together the right people and ideas to produce work that is clear, credible, and effective. This is a hybrid role based in San Francisco, with three days a week in the office. In this role, you will: Develop integrated marketing strategies for business product launches, new capabilities, model releases, and research moments. Partner with Product, Product Marketing, and Research to understand what’s changing, identify the right audiences, and shape positioning and launch narratives. Lead campaigns that bring launches to life across creative, communications, customer marketing, developer marketing, growth, sales, and paid, owned, and earned channels. Identify customer stories, product demonstrations, and creative concepts that make
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: We are seeking an experienced Product Marketing Manager with a strong background in engaging developer audiences and delivering impactful go-to-market programs for native AI and enterprise companies. This role requires someone who is both technically savvy and strategic, with a proven track record of crafting compelling product narratives and building marketing assets that resonate with technical decision-makers. This role is specifically focused on our Model API product offering at Baseten. If you’re passionate about AI infrastructure, developer engagement, and simplifying complex technologies for real-world adoption, we want to hear from you. RESPONSIBILITIES: Positioning & Messaging: Develop clear and differentiated messaging that articulates the value of Baseten’s inference platform to developers and enterprise customers. Narrative Development: Shape how the market thinks about closed-to-open weights models and what matters most when building inference. Go-to-Market Strategy: Own the launch process for new features and products, collaborating closely with product, engineering, sales, and growth teams. Content Development: Create high-quality marketing assets, including white papers, technical blogs, demos, and customer case studies. Sales Enablement: Build resources and programs that empower our sales teams to effectively communicate Baseten’s capabilities and benefits. Market Insights: Understand the
From $164.7K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a Senior Product Manager for Signal Lifecycle within Trust & Safety, you'll own the product strategy for the ML platform that powers how Pinterest trains, evaluates, deploys, and measures content safety models at scale. You'll lead the development of ML Signal Management — making ML signals first-class entities with unified metadata and identity across systems. Partnering deeply with ML engineering, data science, content safety, and enforcement systems, you'll drive a platform whose scope is expanding from T&S into content quality, ads safety, and beyond. What you'll do: Own and drive the Signal Lifecycle product roadmap, including ML Flywheel infrastructure, auto-deployment, model onboarding, golden dataset management, and signal performance measurement Define and ship ML Signal Management — a unified backbone that elevates ML
About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need
About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work
Other cities to consider
More places hiring for this role
Get new model designer jobs in United States by email
Daily job updates · Unsubscribe anytime