About the Team API Enterprise Controls is part of the API Infrastructure organization and owns the platform capabilities that help developers, startups, and enterprises adopt the OpenAI API securely and confidently. We build the systems underneath our APIs and developer platform across authentication and identity, service accounts and key management, secure networking, compliance, auditability, observability, and operational controls. Our users are developers and teams running critical applications on OpenAI, and we partner closely with Product, go-to-market, security, and infrastructure teams to turn their most important needs into reliable, intuitive platform capabilities. About the Role We are looking for an exceptional backend software engineer to help define and ship the enterprise capabilities our API Platform needs to scale.; this is a product-engineering role grounded in deep backend systems. You will work across databases, streaming systems, request routing, authentication, and developer-facing APIs while bringing strong product judgment, developer empathy, and attention to the small details that make a platform easier to understand, trust, and operate. You will lead large cross-functional initiatives, work closely with Product and go-to-market teams, engage directly with sophisticated users, and carry ambiguous needs from discovery through design, launch, and iteration. In this role, you will: Own backend product capabilities end to end across authentication and identity, service accounts and key controls, secure networking, compliance, observability, and operational workflows. Partner with Product, go-to-market, security, infrastructure teams, and sophisticated customers to identify needs, shape the roadmap, and lead large cross-functional projects from design through launch. Design developer-facing APIs, system behavior, configuration, error handling, safe defaults, auditing, and notifications with exceptional care for the details that define a great dev
Jobs in United States
Ai Platform Engineer in United States
5,290 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai platform engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem. About the Role We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives. In this role, you will: Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products. Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration. Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency effi
About the Team API Multimodal builds the developer-facing products and infrastructure that bring OpenAI’s image, audio, and real-time model capabilities into the world. We are responsible for high-scale APIs for image generation, speech transcription, speech generation, and low-latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use customer feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products and distributed systems behind OpenAI’s image, audio, and real-time APIs. You will work across model integration, API design, and production infrastructure to turn new research capabilities into reliable developer experiences. This hands-on role combines backend and systems depth with product judgment: you will own projects end to end, partner with Research, Inference, and Safety, and help make multimodal AI useful at scale. Model training experience is not required. In this role, you will: Design, build, and ship developer-facing APIs and backend services that serve frontier models. Architect low-latency streaming, request, session, and model integration systems that make complex multimodal interactions reliable and intuitive at scale. Work directly with Research to bring new model capabilities into production, shape the systems around them, and incorporate feedback from real-world developers and customers. Own the availability, latency, scalability, and cost efficiency of the services you build. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards. Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed syste
About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including cloud-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re looking for a backend engineer who can quickly understand OpenAI’s models, products, and systems, then adapt first-party deployments for other cloud platforms. You’ll build backend services, APIs, SDK integrations, authentication flows, and cloud service infrastructure that let developers use OpenAI capabilities in the cloud environments where they already build. This role involves working across teams, sometimes embedded with partner product groups, to ship products quickly and across multiple platforms at the same time. It’s a strong fit for engineers who have built developer tools, especially AI-powered tools, communicate clearly across technical boundaries, and can shape architectures that support different deployment models; experience building cloud services is a strong plus. In this role, you will: Build backend and infrastructure systems that extend OpenAI’s API platform into cloud-native environments, like AWS. Design and ship cloud-contained products that allow customers to use OpenAI capabilities while keeping workloads and data within cloud environments. Help stand up cloud-hosted Codex experiences powered by the OpenAI Responses API. Build the infrastructure and runtime abstractions
About the Team OpenAI’s API Platform organization builds the products and infrastructure that help first-party and third-party developers build with OpenAI models. We ship the API primitives, tools, SDKs, documentation, playgrounds, and platform experiences that make OpenAI’s capabilities reliable, understandable, and useful in production. The API Experience team is focused on the end-to-end developer experience for the OpenAI API. We own the surfaces developers touch every day: docs, SDKs, the Playground, examples, onboarding flows, and the systems that help developers go from first request to production deployment quickly and confidently. About the Role We’re looking for full stack and frontend engineers to help define and build the next generation of OpenAI’s developer experience. In this role, you’ll work across frontend product surfaces, backend systems, SDK and documentation pipelines, and API workflows that serve millions of developers and companies. You’ll partner closely with product, design, research, API engineering, and developer-facing teams to make complex AI capabilities simple to understand, easy to test, and safe to launch in real-world applications. This is a highly cross-functional role for someone who cares deeply about craft, developer empathy, reliability, and product velocity. In this role, you will: Build and scale developer-facing products including the OpenAI API Playground, documentation experiences, onboarding flows, examples, and API workflow tools. Own full stack projects end to end, from product definition and UX collaboration through backend implementation, launch, measurement, and iteration. Improve the systems that generate, maintain, and publish SDKs, API references, docs, guides, and developer examples. Partner with API, research, design, and infrastructure teams to bring new model capabilities and API primitives to developers in a clear, usable way. Use developer feedback, product analytics, and direct customer insight to identif
From $186K/yr
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity Are you ready to step into a pivotal leadership role where your engineering depth directly shapes the future of our core platform? As our new Engineering Manager, you will lead a talented, distributed team across US and EU time zones, acting as the critical manager bridging regional collaboration. Our Cloud Foundation team is the backbone of the New Relic platform. In this role, you won't just manage tasks; you will mentor and empower engineers, transitioning our operational framework from a reactive state to a culture of proactive ownership and engineering excellence. You will oversee critical global initiatives, including major regional expansions into FedRAMP High / IL4, India, and Australia. If you thrive on solving complex multi-cloud challenges at an exabyte scale while helping engineers grow in their careers, this is your opportunity to make a lasting impact. What you'll do Empower & Mentor: Lead and nurture a high-performing engineering team across the US and EU, facilitating career development, performance growth, and a collaborative team culture. Drive Strategic Ownership: Champion a shift from reactive delivery to proactive technical ownership, establishing best practices for platform reliability and cross-regional alignment. Lead Regional Expansions: Architect and execute key global infrastructure expansions across complex environments (including FedRAMP High / IL4, India, and Australia). Architect for Extreme Scale: Guide decisions around micr
From $399.4K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Content Platform team at Roblox powers the infrastructure behind every asset used across the Roblox ecosystem—enabling creators and developers to bring their visions to life at global scale. From 3D models and images to videos and audio, our platform manages the complete lifecycle of all assets essential for immersive experiences, supporting one of the largest services in the world at over 100+ million requests per second. Our mission is to deliver a seamless, reliable, and innovative content system that empowers creators, supports record-breaking games, and ensures the highest standards of performance and safety for our community. As the Technical Director for Content Platform, you will lead multidisciplinary engineering teams responsible for the technical and product vision of Roblox’s asset infrastructure. You will own the lifecycle of every asset—from creation and upload to storage, indexing, delivery, and rendering in the game client. Your leadership will be critical in scaling our systems, optimizing distributed infrastructure, and enabling new possibilities for creators and players alike. You Will: Define and drive the long-term strategy, architecture, and priorities for the Cont
From $152.8K/yr
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As Engineering Manager for the Agent Execution group , you'll lead more than 10 engineers who build foundations for AI agents to run safely and effectively across GitLab. You'll give the group clarity, remove obstacles, and help teams move quickly while protecting what matters. You'll contribute to design reviews on sandboxing and agent observability, work alongside staff engineers, and use modern AI coding tools and agent harnesses in your work. What you’ll do Lead the Agent Tools, Agent Observability, and Runner Execution teams across frontend, backend, and AI engineering, with each team anchored by a s
About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for an AI Agent Security Architect to function as the primary technical authority for Replit’s autonomous and AI agent security blueprint. In this critical role, you will design, implement, and maintain the runtime defense systems, guardrail frameworks, and sandboxing architectures that govern AI agents executing code, invoking tools, and reasoning across our platform. You will be a key technical contributor—leading high-impact AI security initiatives and bridging the gap between non-deterministic AI behavior and rigorous cybersecurity controls for both engineering and executive leadership. What You'll Do AI Agent Security Strategy & Technical Execution AI Agent Security Blueprint: Define the long-term vision and architectural patterns for securing autonomous agent workflows, Model Context Protocol (MCP) integrations, multi-turn reasoning loops, and multi-agent coordination. Runtime Guardrails & Policy Enforcement: Architect and deploy dynamic input/output guardrail systems, semantic firewalls, and real-time intent verification filters to prevent goal hijacking, system prompt leaks, and indirect prompt injections. Agent Execution & Tool Sandboxing: Partner with Infrastructure and AppSec teams to design secure, short-lived, micro-isolated environments (e.g., microVMs, WebAssembly, container sandboxes) where agents can dynamically execute code, run shell commands, and interact with host operating systems safely. Agentic Threat Modeling & Red Teaming: Conduct specialized threat modeling against non-deterministic systems. Lead automated and manual AI red-teaming initiatives to uncover vulnerabilities in RAG context pipelines, vector stores, and tool-calling interfaces. Identity
About the Team The Plugin Ecosystem team builds the platform and product experiences that let people extend ChatGPT and Codex. We work on plugins, skills, connectors, interactive apps, and open standards like the Model Context Protocol (MCP). We make plugins easy to discover, install, and use, ensure they’re invoked at the right time, and help people find new ways to get value from them. We want anyone to be able to turn a useful workflow into a plugin, share it, and have other people use it. A plugin can package instructions and skills with connections to the tools and data it needs. Our work spans creation and publishing, reliable execution across our products, clear permissions and approvals, and the controls admins need to bring plugins to their organizations. We work closely with research to improve plugin quality as models evolve. About the Role We’re looking for product-minded engineers to build the systems behind plugins and improve how models use them. Depending on your focus, you may scale generalist infrastructure and identity-related integrations across products, or improve plugin quality at the intersection of backend engineering and applied AI or work on the product experience itself to drive plugin usage. You’ll work across teams and own problems from diagnosis and design through implementation and release. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and ship APIs, SDKs, and services that developers use to extend ChatGPT and Codex. Build intuitive experiences that help users discover, install, and use plugins to get more done. Make plugins easier to create, test, publish, update, and share. Improve when and how models use plugins, from choosing the right plugin to completing a task. Work with Research to diagnose failures and measure improvements as models evolve. Improve plugin reliability and interaction quality acros
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. The AI Studio is the team responsible for designing, building, and operationalizing internal software platforms across every functional department of the company, including People Operations, Customer Support, Sales, Marketing, Finance, Recruiting, and Executive Operations. We treat each internal department as its own product surface, and in truth as its own startup, with its own customers, its own metrics, and its own reasons for existing. We ship custom software, built on Replit’s own platform, to solve operational problems that off-the-shelf SaaS tools cannot serve well at our scale and pace of growth. Position Summary The Member of Technical Staff (AI Builder) is a full-time builder role with substantial autonomy. You will design, architect, and ship AI-powered internal platforms that run the operations of the company across many departments at once. This is not a role for someone who wants a narrow, well-defined lane. It is a role for someone who wants to walk into an ambiguous business problem, understand it well enough to argue about it with the department lead, and then ship the software that fixes it. We are looking for people who genuinely understand the mechanics of a business: how revenue is actually made, why recruiting velocity matters, what makes support scale or break, how finance closes a month, and why each department is critical to whether the company wins or loses. You do not need to have run every function, but you need to respect why each one exists and be able to reason about it like an operator, not just an engineer. Because we treat each department as its own startup, we strongly favor people who have either run their own business or worked on a fast-scaling startup. You know what it feels like
About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in AWS-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including AWS-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re hiring Machine Learning Engineers to build and improve the AI systems that help strategic partners adapt OpenAI models to important use cases in cloud-native environments. This role spans post-training workflows, evaluation, data pipelines, model behavior, and API/infrastructure integration. You’ll work at the boundary between partner needs and core ML systems: helping teams understand what is and isn’t working, diagnosing issues in training and evaluation workflows, and turning those learnings into improvements to the underlying platform. You should enjoy working with external technical partners, extracting the real goal from messy requests, and pushing back or reframing when the requested experiment is not the highest-leverage path. You’ll collaborate closely with Research, Applied, Safety Systems, infrastructure teams, and external technical partners to solve ambiguous model-performance problems. When you succeed, strategic partners and internal teams will be able to improve model behavior with confidence, driving measurable product improvements while the systems behind that work become more reliable, scalable, and effective over time. In this role, you will Partner with strategic customers and in
$293K – $325K/yr
About the Team The Statsig team at OpenAI builds and operates the experimentation platform that powers product development, measurement, and decision-making across the company. We partner closely with product, engineering, and infrastructure teams to ensure experiments are trustworthy, statistically rigorous, and scalable to the needs of frontier AI products. Our mission is to help teams make better decisions through reliable experimentation. We care deeply about statistical correctness, pragmatic solutions, and building systems that researchers and engineers can trust at massive scale. The team operates at the intersection of experimentation methodology, data infrastructure, causal inference, and product analytics. We are looking for experienced experimentation experts who want to shape the future of experimentation in the AI era. About the role: We're seeking a Data Engineer to take the lead in building our data pipelines and core tables for OpenAI. These pipelines are crucial for powering analyses, safety systems that guide business decisions, product growth, and prevent bad actors. If you're passionate about working with data and are eager to create solutions with significant impact, we'd love to hear from you. This role also provides the opportunity to collaborate closely with the researchers behind ChatGPT and help them train new models to deliver to users. As we continue our rapid growth, we value data-driven insights, and your contributions will play a pivotal role in our trajectory. Join us in shaping the future of OpenAI! In this role, you will: Design, build and manage our data pipelines, ensuring all user event data is seamlessly integrated into our data warehouse. Develop canonical datasets to track key product metrics including user growth, engagement, and revenue. Work collaboratively with various teams, including, Infrastructure, Data Science, Product, Marketing, Finance, and Research to understand their data needs and provide solutions. Implement ro
From $109K/yr
The Agent Research and Tooling team, part of MongoDB's AI Builder Experience organization, owns the platform layer around agents: how teams author, distribute, evaluate, monitor, and improve agent skills and agent behavior. We are hiring a software engineer to build and maintain the tooling, evaluation systems, and quality gates behind MongoDB's agent skills. This is a software engineering role at the intersection of developer tooling, applied AI, and software quality. You will take loosely defined agent and tooling problems, break them into workable plans, and ship durable internal systems: command-line tools, reusable libraries, evaluation harnesses, and CI workflows. This role is open to remote work in the US or can be based out of any of our US offices. What you'll do Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI Integrate tooling into GitHub Actions and other CI workflows, including secrets, annotations, exit codes, and artifacts Build code-generation quality checks, such as anti-pattern catalogs and linting for AI-generated MongoDB code Investigate real failures such as nondeterministic results, false positives, and unsafe generated guidance, and turn them into reusable improvements Collaborate with engineers, security partners, and product teams; communicate trade-offs, risks, and ownership across teams Examples of the problems you'll solve How can tests verify an agent tool's
Other cities to consider
More places hiring for this role
Get new ai platform engineer jobs in United States by email
Daily job updates · Unsubscribe anytime