Jobs in United States

Design Quality Engineer in United States

2,264 active opportunities · Updated October 2026

Explore current design quality engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Intelligence and Investigations team is dedicated to ensuring the safe, responsible deployment of AI by rapidly detecting and mitigating abuse. Our team leverages the latest testing methodologies to uncover vulnerabilities and emerging threats, helping safeguard OpenAI’s products and users. We work closely with cross-functional partners across product, policy, and engineering to drive a comprehensive defense strategy against evolving adversarial challenges. About the Role As a Red Team Specialist focused on cyber, you will help answer two practical questions: What cyber capabilities can our models provide to real-world attackers, and do our safeguards remain effective when those attackers use increasingly sophisticated techniques? The role combines scaled evaluation with expert-driven testing. You may bring deeper experience in cybersecurity and use that expertise to judge whether a model’s behavior meaningfully changes attacker capability. Alternatively, you may bring deeper experience in model evaluations, automation, or agentic harnesses and apply those skills to building rigorous cyber testing. We do not expect every candidate to be equally deep in both areas, but successful candidates will have a strong foundation in one and enough fluency in the other to work effectively across the boundary. Most of your work will focus on model cyber capabilities and safeguards; you will also spend a portion of your time testing novel abuse risks in agentic systems. This role is located in San Francisco, CA or Seattle, WA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and run rigorous evaluations of model cyber capabilities and safeguards, including policy adherence, correct refusal, over refusal, and resilience to jailbreaking and other adversarial techniques. Conduct hands-on testing to understand what models can enable when used by experienced security practiti

AWSRestAIGo
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed Engineers on our engineering team who want to work at the intersection of deep infrastructure work and direct customer impact. As an FDE, you'll partner with leading AI companies and foundation labs on cloud architecture, networking, storage, containerization, sandboxing, and more — helping them design and ship production infrastructure on Modal's platform. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the infrastructure stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and deploy massive-scale production workloads on Modal Lead technical discovery and architect

AWSAzureGCPDocker
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and proactive Security Engineer to help us build, maintain, and continuously improve the security posture of our rapidly growing ML infrastructure platform. As one of the first dedicated security hires at Baseten, you will work cross-functionally with engineering and operations teams to ensure we’re meeting the highest standards of confidentiality, integrity, and availability. You’ll have an opportunity to shape our security strategy and best practices from the ground up, influencing the way our platform handles sensitive data for both internal and external stakeholders. RESPONSIBILITIES Security architecture and design: Collaborate with engineering teams to design and implement secure systems and infrastructure, including cloud (AWS/GCP) environments and container orchestration platforms. Vulnerability management: Lead proactive vulnerability assessments, pen tests, and remediation efforts to ensure our products and infrastructure remain secure. Incident response: Develop and maintain incident response processes, including detection, analysis, containment, eradication, and post-incident reviews. Identity and access management (IAM): Oversee IAM strategies and tools to ensure the right people have the right level of access to our systems and data. Security compliance and audits: Work closely with operations to ensure compliance with relevant standards (e.g., SOC 2, ISO 27001) and

AWSGCPCI/CDMachine Learning
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Infrastructure Security Engineer to design and secure the core systems that power our platform. This role focuses on building security directly into our infrastructure—from container isolation and orchestration to identity and secrets management in a multi-tenant, cloud-native environment. You’ll work closely with engineering teams to define secure primitives and ensure our platform is resilient, scalable, and trustworthy by design. This is a hands-on, deeply technical role focused on real systems, not compliance or policy. What You'll Do: Platform & Runtime Security Design and improve isolation mechanisms for multi-tenant workloads (containers, sandboxing, execution environments) Strengthen boundaries between customers, workloads, and internal systems Identify and mitigate risks in distributed, dynamic compute environments Container &

AWSGCPKubernetesAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Product Engineer on the Dedicated Inference team, you'll shape the state-of-the-art developer experience for deploying and operating AI workloads in production. From the CLI and SDKs to APIs, observability, and debugging workflows, you'll build the tools customers rely on every day to manage mission-critical inference deployments. Few teams at Baseten have as much breadth and visibility as Dedicated Inference. The team is often at the forefront of new product development, giving engineers the opportunity to shape the experience of some of our most important customers. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Dedicated Inference team: Chains for multi-component workflows Asynchronous inference Model APIs for frontier models Model training built for production inference RESPONSIBILITIES Implement new features and products for the team Design ergonomic APIs and abstractions to solve customer problems Fix bugs and resolve customer issues with urgency Work across the stack - regardless of where you start, you’ll end up touching both React Components and Kubernetes Pods Work closely with the product and forward deployed engineering teams to develop and drive new product ideas REQUIREMENTS Bachelor's degree or higher in Computer Science or related field Proficient coding abilities in one or more popular programming or scripting languages; Python, Go, or Javascript proficie

JavaScriptPythonJavaReact
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues. Apply and scale optimization techniques across a wide range of ML models, particularly large language models. Collaborate with a diverse team to design and implement innovative solutions. Own projects from idea to production. REQUIREMENTS Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field. Experience with one

PythonDockerKubernetesRest
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ROLE We're looking for a Solutions Product Marketing Manager to define and own Baseten's up-market, enterprise, and industries go-to-market strategy. You'll build the playbook that helps Baseten win large, complex deals and the vertical go-to-market strategies that let us compete in Financial Services, and Health & Life Sciences, and the use cases that scale across all of them. This is a highly cross-functional, largely greenfield role. You'll partner closely with the Industries sales team, product marketing, demand gen, and executive leadership to build on our enterprise go to market motion. RESPONSIBILITIES Enterprise Deal Playbook & Buyer Enablement Build the enterprise deal playbook and full bill of materials (BOM) — pricing guides, one-pagers, TCO/ROI guides, landing pages, and event/field collateral — mapped to every stage of the buyer's journey Build the system for keeping every sales asset current — clear ownership, last-updated dates, and a refresh cadence Develop industry-specific event concepts featuring anchor customers and AI-native accounts to build credibility with target executive audiences Vertical GTM & Industry Go-to-Market Kits Build vertical sales kits and ICP/positioning frameworks for priority industries Own industry-specific landing pages, messaging, and TCO guides, extrapolating from early AI-native accounts to larger enterprise targets Design and help launch an ABM experiments in p

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ROLE We’re looking for a high-performing strategic finance professional to join our growing GTM Finance team. Our business grows with our customers' usage, which makes the finance function highly strategic at Baseten: growth, pricing, margin, and capacity decisions are business model decisions. You'll sit at the center of them, partnering directly with GTM leadership and reporting into a finance team with a seat at the table for the calls that shape the company's trajectory. This role is ideal for someone with 3 to 7 years of experience across strategic finance, investing, and/or investment banking who wants broad exposure to company-building inside a fast-scaling AI infrastructure company. Experience at a usage-based software company is a plus. RESPONSIBILITIES Own financial planning, forecasting, and budgeting processes for the GTM org Build and maintain financial models across revenue, S&M spend, headcount, and strategic bets Analyze the metrics that define a usage-based business – ARR, gross margin, consumption trends, retention, and GTM efficiency Partner with GTM leaders to set targets, evaluate growth initiatives, shape pricing, and design sales compensation Help prepare board materials, investor updates, and fundraising analyses Improve financial reporting, dashboards, and operational rigor so our infrastructure scales as fast as our revenue Work cross-functionally to turn ambiguous business questions into

SQLRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ABOUT THE ROLE Baseten is building the infrastructure layer for AI — and we're now building the team that will scale how we take it to market. This is a foundational hire on our GTM Strategy & Revenue Operations team, sitting at the intersection of strategic planning and field execution. You'll work directly with the CRO, Head of Revenue Operations and partner closely with Finance, field leadership, and our Central Ops team. On any given week, you might be refining our pipeline generation model, building a territory coverage analysis, designing a new GTM motion, or partnering with a regional leader to understand what's driving a trend in their pipeline. This role requires someone who can think rigorously, build things from scratch, and operate with speed and judgment in an environment where the playbook is still being written. This is a rare opportunity to be an early GTM strategy and operations hire at one of the fastest-growing companies in AI infrastructure — and to help define how we scale. WHAT YOU'LL DO: GTM Planning, Target Setting & Market Intelligence Contribute to the annual and quarterly GTM planning process in partnership with Finance, including headcount modeling, ramp assumptions, and revenue target-setting Build and maintain coverage models aligned to 2-year growth projections, incorporating territory design, account segmentation, and capacity planning Support quota framework development, helping

SQLMachine LearningAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Customer Marketing Manager who can own the full customer evidence motion at Baseten: building the systems that capture customer stories, running the co-marketing programs that amplify them, and developing the channels and assets that get those stories in front of the right people. Our customers are ML engineers and AI teams deploying serious workloads — and the stories they tell about what they've built matter. We've earned trust with some of the most demanding technical teams in the industry, and this role exists to turn that trust into evidence. RESPONSIBILITIES Co-Marketing Execution Serve as the DRI for every customer co-marketing launch end to end — managing timelines, coordinating internal and external stakeholders, and driving the process from first outreach to final publication Own the single source of truth for what's in flight across all customer co-marketing activity Coordinate with design, social, and sales to ensure every asset is built, approved, and distributed correctly Customer Evidence & Asset Library Own the customer evidence library: written case studies, video stories, customer quote repository, logo library, and sales snippets ensuring all assets stay current and are tagged and accessible for sales and marketing use Run the monthly operating rhythm: new logo additions from closed-won opportunities, asset updates, and customer health checks Programs & Channels I

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b

PythonKubernetesLinuxMachine Learning
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications. You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Core Engineering Responsibilities Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap Performance & Innovation Impl

AWSMachine LearningAIC++
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As a senior member of Baseten's Platform Team, you will own the systems that let every engineer at Baseten prove their code works before it reaches production. Our product runs mission-critical AI inference for customers who measure downtime in dollars per second, which means our internal bar for correctness, performance, and failure tolerance has to be exceptional. Your focus is the full testing stack: fast and reliable unit test tooling, integration harnesses that spin up realistic environments on demand, load and performance testing for GPU-backed inference workloads, and resilience testing that deliberately breaks things so our customers never have to find out what happens when a node dies mid-request. This is a builder role with org-wide leverage. You won't be writing tests for other teams — you'll be building the frameworks, harnesses, and feedback loops that make writing good tests the path of least resistance, and you'll set the standards for what "well-tested" means at Baseten. RESPONSIBILITIES Own Baseten's testing strategy end to end — define the standards, the tiers, and the tooling that engineering teams build against. Build and maintain unit, integration, load and performance testing frameworks Design end to end test infrastructure that provisions realistic dependencies

PythonDockerKubernetesCI/CD
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As an early member of Baseten's Platform Team, you will be pivotal in building internal infrastructure to support our engineering organization. You will own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across Baseten to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. If you are passionate about elegant solutions—like streamlined monorepos, lightning-fast CI pipelines, and thoughtfully designed shared libraries—you'll thrive at Baseten. RESPONSIBILITIES Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces,

PythonKubernetesGitRest
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming the internet’s next interface, but a production-grade Voice AI system is "hard to build" . You’ll join a small founding team of Baseten Voice AI, focused on bringing state-of-the-art open source models into production for Voice AI customers across productivity, customer service, clinical conversation, creator tools, education, and more. You’ll make a meaningful impact on people’s daily lives and help reshape these industries. This is a high-impact, high-ownership role. You will be the primary owner of Baseten Voice AI - our in-house inference stack to power Voice AI models - from product roadmap through engineering implementation. You’ll partner closely with Forward Deployed Engineers, Model Performance Engineers, and sister engineering teams to push the boundaries of Voice AI. EXAMPLE INITIATIVES: Develop world-class model serving stack for state-of-the-art open-source voice models - reduce end-to-end and tail latency (p95/p99), increase throughput, and improve GPU efficiency via profiling, runtime tuning, and server-level optimizations. Build large-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS, and agent components with streaming I/O to meet customer SLOs. Design tight training and inference iteration loops for voice model customization - enable fast evaluation, safe rollout, and rapid experimentation for custom voice model development. Past projects:

PythonDockerKubernetesRest
🔔

Get new design quality engineer jobs in United States by email

Daily job updates · Unsubscribe anytime