Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. You may
Jobs in United States
Audio Inference Engineer in United States
266 active opportunities · Updated October 2026
Showing
15 jobs
Explore current audio inference engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed ML Engineers who want to work at the intersection of deep technical work and direct customer impact. As an ML FDE, you'll partner with leading AI companies and foundation model labs to help them achieve state-of-the-art performance on their most demanding workloads — LLM serving, model training (SFT, RLHF), audio pipelines, scientific computing, and more. You're helping teams reach outcomes most engineers can't on their own. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the AI stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and optimize production AI workloads
About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required. In this role, you will: Design, build, and ship developer-facing APIs and backend services that serve frontier models. Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale. Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers. Own the availability, latency, scalability, and cost efficiency of the services you build. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards. Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed syste
About the Team The Applied team at OpenAI safely brings cutting-edge technology to the world. We have released widely used products such as ChatGPT, Sora, and the OpenAI API, powering models including GPT-5 and a growing set of multimodal capabilities across text, image, audio, and video. Our team also manages large-scale inference and platform infrastructure that supports these experiences at global scale. With much more on the horizon, our impact continues to grow. Our customers build fast-growing businesses using our APIs, unlocking product capabilities that were previously unimaginable. ChatGPT and Sora exemplify the breadth of what’s now possible across text, image, audio, and video experiences. As these capabilities expand, we prioritize the responsible use of our technology, emphasizing safe and thoughtful deployment over unchecked growth. Within Applied Engineering, the Ads Monetization team in Financial Engineering builds the core systems dealing with all the money flows for ChatGPT Ads. These systems are a combination of low-latency, high scale, high reliability, while being built in a financially correct, accurate, auditable and explainable way. This role sits at the intersection of ads delivery, data engineering, and financial systems. In this role, you will: Architect and build the core monetization systems for ChatGPT Ads. Build and operate the core services and pipelines that power ads monetization end-to-end, from event capture and validation through aggregation, pricing, metering, and ultimately producing billable outputs. Define and implement the source of truth for ads monetization data, including schemas, data models, and invariants that ensure outputs are consistent, explainable, and auditable. Own correctness and reconciliation: align production outputs with downstream invoicing/finance requirements, build controls/monitors, and close gaps through investigations and backfills. Develop across the stack to create comprehensive billing integration
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten's GTM org is in hyper growth. As it grows and matures, the needs of the GTM stack get more sophisticated with scale — and this team exists to stay ahead of those needs. Our GTM Engineering team plays a critical role in building and maintaining the connective tissue across Baseten’s GTM tools, processes, and user experience for the field. The GTM tooling landscape is changing fast, and the teams that win are the ones that adapt and iterate the fastest. This role exists to make sure Baseten is one of them. You'll design, build, and ship AI-powered workflows that scale our GTM functions as a competitive advantage. We want someone who can walk in, audit what we have, identify what we're missing, and start shipping fast. You know when to reach for Clay and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader systems architecture tomorrow. And you bring a point of view — on our stack, on what we should be building, and on where AI can do something low-code tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for the field — build the agents and automations that give reps and managers real leverage, off-loading the manual and repetitive work. Reach for AI where it does something low-code can't. Get insights in front of reps — turn Salesforce, warehouse, and usage data into the dashboards, scores, and alerts reps
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and detail-oriented GRC (Governance, Risk, and Compliance) Manager to build, support, and continuously enhance Baseten’s security governance, compliance, and privacy programs. As one of the early members of our security organization, you will play a key role in ensuring our platform meets and exceeds the highest standards for privacy, trust, and regulatory compliance. In this role, you’ll work cross-functionally with engineering, operations, legal, and leadership teams to develop policies, manage audits, and implement controls aligned with frameworks such as SOC 2, ISO 27001, ISO 27701, and FedRAMP. You’ll be instrumental in building scalable processes to manage risk, support customer assurance, and uphold Baseten’s commitment to security and compliance as we grow. RESPONSIBILITIES Governance & Policy Development: Design, implement, and maintain security governance frameworks, policies, and procedures that align with Baseten’s risk posture and industry best practices. Risk Management: Build and manage the company-wide risk assessment program, identifying, tracking, and mitigating key security and compliance risks. Compliance Operations: Lead efforts to achieve and maintain compliance with SOC 2, ISO 27001/27701, HIPAA, FedRAMP and other applicable standards and regulations. Audit & Certification Management: Coordinate external audits and certification processes, ensuring e
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Revenue Accounting Manager to own billing, contract review, and revenue recognition as Baseten scales. This is a hands-on, individual-contributor role for someone who wants full ownership of the revenue cycle at a company where deal structures are getting more complex: usage-based pricing, committed capacity, credits, multi-year enterprise contracts, and new business models coming online. You'll join Lauren, who built Baseten's quote-to-cash function from the ground up, to take on billing, contract review, and revenue recognition as deal volume and complexity grow. Having completed our first year-end audit, we're now focused on tightening contract review processes, close procedures, and reporting rigor to support the scale ahead. You'll partner closely with Lauren, FP&A, Sales, Legal, and Revenue Operations to make sure every deal is structured, billed, and recognized correctly from day one. Baseten is building the infrastructure layer for AI-native companies, and we're scaling quickly - in deal volume, contract complexity, and customer size. If you want real ownership over a growing function on a lean team, this role offers real scope. RESPONSIBILITIES Revenue Recognition and Technical Accounting Own revenue recognition under ASC 606 across all contract types, including usage-based/consumption arrangements, committed capacity deals, credits, multi-year contracts, and new business model
About the Team The OpenAI Audit Team is on a mission to build the future of internal audit from the ground up. Our ambition will be powered by a high-energy, technically exceptional team with the judgment, intellectual curiosity, and creativity to harness the latest advances in AI and design a truly next-generation audit function. As AI reshapes how work is performed across the enterprise, Internal Audit will use AI, automation, and data analytics to identify and assess the most significant and emerging risks across the business, including technology, cybersecurity, finance, compliance, operations, and data. We will build trusted partnerships at every level—from the Board of Directors and senior leadership to the teams delivering on OpenAI’s mission every day. We will operate as both an independent assurance provider and a trusted advisor, bringing an objective and pragmatic perspective to critical decisions. By engaging closely with management while preserving our independence, we will help the business innovate responsibly, move with confidence, and manage risk without creating unnecessary barriers. About the Role As the Cybersecurity & Technology Audit Leader, you will help shape the strategy, methodology, technology, and culture of a new audit function. You will lead complex audits and advisory reviews of cybersecurity, technology, data, and AI-related risks while advising leaders on practical ways to strengthen governance, resilience, and risk management. You are an experienced, hands-on professional who combines deep technical expertise with strong business judgment. You can quickly move between executive-level governance questions and detailed technical analysis, use data to identify and assess risk, and form clear, well-supported conclusions in fast-moving or ambiguous situations. You will have meaningful influence over how the function develops, including how we apply AI, automation, and analytics to audit planning, testing, monitoring, and reporting. T
The Regulatory Risk Group Manager is accountable for management of complex/critical/large professional disciplinary areas. Leads and directs a team of professionals. Requires a comprehensive understanding of multiple areas within a function and how they interact in order to achieve the objectives of the function. Responsible for driving customer remediation analytics to completion including execution and validation. This position will be part of the USCC Centralized Customer Remediation Team (CCRT). The objective of the USCC CCRT is to quickly assess and identify customer impacts and ensure remediation is completed promptly with the appropriate detailed documentation required by Internal Audit and the Regulators. These activities will be conducted in close partnership with the Issue Owners and Managers within U.S. Consumer Cards. Excellent communication skills are required to negotiate internally, often at a senior level. Highly developed communication and diplomacy skills are required, to guide, influence, and convince others, in particular colleagues in other areas and occasional external customers. This role is a people manager role with full management responsibility of a team or multiple teams, including management of people, budget and planning, to include performance evaluation, compensation, hiring, disciplinary actions and terminations and budget approval. Responsibilities: Manage a team of remediation analysts and associated book of work for the team. Perform training, reviews, coaching, and mentoring of staff. Over-sees, designs, and manages client remediation population and impact analyses for USCC owned and impacted issues, including those that are Regulatory. Provide expertise and guidance related to the USCC remediation processes and policies, partners with other department leaders to define, prioritize, and execute remediation programs. Create comprehensive documentation that confo
The Regulatory Risk Officer is a strategic professional who stays abreast of developments within own field and contributes to directional strategy by considering their application in own job and the business. Recognized technical authority for an area within the business. There are typically multiple people within the business that provide the same level of subject matter expertise. Responsible for driving customer remediation analytics to completion including execution and validation. This position will be part of the USCC Centralized Customer Remediation Team (CCRT). The objective of the USCC CCRT is to quickly assess and identify customer impacts and ensure remediation is completed promptly with the appropriate detailed documentation required by Internal Audit and the Regulators. These activities will be conducted in close partnership with the Issue Owners and Managers within U.S. Consumer Cards. Excellent communication skills are required to negotiate internally, often at a senior level. Highly developed communication and diplomacy skills are required, to guide, influence, and convince others, in particular colleagues in other areas and occasional external customers. This role is for an individual contributor. Responsibilities: Subject matter expert related to data analytics and remediation methodology. May train, mentor, and coach lower-level staff. Owns, designs, and manages client remediation population and impact analyses for USCC owned and impacted issues, including those that are Regulatory. Provide expertise and guidance related to the USCC remediation processes and policies, partners with other department leaders to define, prioritize, and execute remediation programs. Create comprehensive documentation that conforms to Internal Audit and Regulatory expectations. Provide thought leadership and effective challenge throughout the data remediation lifecycle (including root
The Regulatory Risk Officer is a strategic professional who stays abreast of developments within own field and contributes to directional strategy by considering their application in own job and the business. Recognized technical authority for an area within the business. There are typically multiple people within the business that provide the same level of subject matter expertise. Responsible for driving customer remediation analytics to completion including execution and validation. This position will be part of the USCC Centralized Customer Remediation Team (CCRT). The objective of the USCC CCRT is to quickly assess and identify customer impacts and ensure remediation is completed promptly with the appropriate detailed documentation required by Internal Audit and the Regulators. These activities will be conducted in close partnership with the Issue Owners and Managers within U.S. Consumer Cards. Excellent communication skills are required to negotiate internally, often at a senior level. Highly developed communication and diplomacy skills are required, to guide, influence, and convince others, in particular colleagues in other areas and occasional external customers. This role is for an individual contributor. Responsibilities: Subject matter expert related to data analytics and remediation methodology. May train, mentor, and coach lower-level staff. Owns, designs, and manages client remediation population and impact analyses for USCC owned and impacted issues, including those that are Regulatory. Provide expertise and guidance related to the USCC remediation processes and policies, partners with other department leaders to define, prioritize, and execute remediation programs. Create comprehensive documentation that conforms to Internal Audit and Regulatory expectations. Provide thought leadership and effective challenge throughout the data remediation lifecycle (including root
About the Team OpenAI Finance helps ensure the organization is set up for success in pursuit of its mission of ensuring that artificial general intelligence benefits all of humanity. Within Finance, the Technical Revenue team partners with Product, Legal, GTM, Strategic Finance, Revenue Accounting, Finance Systems, and the company’s GTM Deal Desk organization to enable scalable, audit-ready monetization. We advise on commercial and contract design, establish defensible accounting positions under ASC 606, and provide clear conclusions and handoffs for downstream execution by Revenue Accounting Operations. About the Role We’re looking for a Senior Manager of Technical Revenue Enablement to lead Revenue Recognition deal advisory for OpenAI’s general B2B commercial activity. Reporting to the Head of Technical Revenue, you will serve as the primary Revenue Recognition counterpart to OpenAI’s Deal Desk organization, partnering with Legal, GTM, Product, Pricing, and Strategic Finance throughout the presignature deal lifecycle. You will own intake and triage, advise on non-standard terms, approve arrangements supported by established policy and precedent, and route novel or strategic matters to the appropriate technical owner. You will build a scalable deal-advisory model by translating ASC 606 into practical contract guardrails, approved language, decision trees, review thresholds, and clear service levels. This role requires deep technical accounting judgment, commercial fluency, strong stakeholder influence, and the ability to make timely, risk-based decisions in a fast-moving environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead the Revenue Recognition deal-advisory function for general B2B transactions, serving as the primary finance counterpart to OpenAI’s GTM Deal Desk organization. Partner early with GTM Deal Desk, Legal, GTM, Pr
About the Team OpenAI Finance is responsible for ensuring the organization is set up for success in pursuit of its mission. The Technical Accounting, Accounting Policy, and Financial Reporting organization partners across Finance to assess complex corporate matters and develop well-supported U.S. GAAP conclusions. The team supports non-routine business activities requiring thoughtful technical judgment, scalable accounting policies, financial statement disclosures, processes, and controls. About the Role You will serve as a senior technical accounting leader and strategic advisor to Finance and business leadership. You will oversee a broad portfolio of complex accounting matters and strategic transactions, including mergers and acquisitions, major commercial arrangements, financing transactions, investments and financial instruments, leases, intercompany transactions and consolidation, impairment, and other matters requiring significant professional judgment and practical implementation. You will advise and influence stakeholders across Corporate Accounting, Financial Reporting, Treasury, Legal, Tax, Strategic Finance, FP&A, Equity, Internal Controls, valuation specialists, and external auditors. You will drive timely resolution of complex accounting questions, align leaders on material judgments and risks, and ensure conclusions are translated into audit-ready memoranda, close entries, financial statement disclosures, and durable controls. You will also help strengthen external-reporting processes, as applicable. This is a highly visible role for an experienced accounting leader who can independently set direction for complex workstreams, advise senior decision-makers, anticipate and escalate material judgments, align cross-functional stakeholders, and carry issues from authoritative research through clear recommendations and disciplined implementation. Location and work model: This role is based in San Francisco and follows a 3 day hybrid work model. In this r
Other cities to consider
More places hiring for this role
Get new audio inference engineer jobs in United States by email
Daily job updates · Unsubscribe anytime