About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform
Jobs in United States
Ai Platform And Agentic Engineer in United States
5,418 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai platform and agentic engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent engineering manager to own and deliver an end to end manageability stack for Data Center Systems. We are seeking an experienced manager who is deeply technical, hands-on, and has a wide system view. You will manage a team of experts, design & build OpenBMC based manageability software stack for NVIDIA’s next generation Data Center Compute Systems. We want to grow our teams with the smartest people in the world. If you're creative and autonomous, we want to hear from you! What you’ll be doing: Own and deliver OpenBMC based manageability stack for next generation Data Center Compute Systems. Own firmware delivered to data centers in terms of quality, reliability and telemetry performance. Manage and lead a distributed team of software engineers to deliver firmware stack with high quality. Work with data center architects and cloud customers for correct requirements and scope implementation to ensure speed of light product development. Work closely with cross functional teams to ensure scalable manageability architecture for all data centers products Drive efficiency, reliability and optimization in firmware architecture from a data center view point. Work closely with customers and internal teams to resolve issues at Speed of Light. What we need to see: BS, MS, or PhD in EE/CS or related field o
At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. About the Role Every consumer experience Mindbody ships is built on top of our consumer platform. The platform is the shared foundation that lets our front-end squads deliver reliable, consistent wellness experiences without rebuilding the plumbing every time. This role exists to answer a deceptively simple question: what does the consumer platform need to become so that front-end squads can move fast on solid ground — and how do we get there from a legacy monolith without breaking the surfaces that depend on it today? This is a senior individual-contributor role for a product leader who is equally comfortable in two rooms: designing API contracts and service boundaries with a senior backend engineering team, and making the cross-surface platform calls that keep multiple consumer squads pulling in the same direction. You'll own the consumer platform roadmap, partner closely with engineering and your downstream partner squads, and be measured on how much faster and more reliably the rest of consumer can build because of the foundation you've shaped. What you'll do Own the consumer platform strategy and roadmap. Define the capabilities the platform must provide so consumer s
From $196.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Engineer on the Detection and Response (D&R) team at Roblox, you’ll protect our user community alongside the underlying platform infrastructure. You’ll design high-fidelity detections, engineer security data platforms, and respond alongside the team during incidents. This is a hybrid in-office role in San Mateo. You Will: Deliver robust D&R capabilities: Engineer high-fidelity detections end-to-end. Lead partners through threat modeling and logging, to deploying actionable alerts, while keeping false positives low. Build security data pipelines: Develop security data pipelines and actively contribute to internal software and data platforms, collaborating across engineering teams. Ensure service reliability: Participate in an on-call rotation to keep detection and response services healthy. Embody security culture: Serve as a trusted security partner across Roblox, helping protect our community and enterprise while fostering a culture grounded in trust, ownership, and shared responsibility. You Have: 3+ years of experience in Security Data Engineering: You have built services that are efficient, reliable, and scalable using programming languages like Golang or Py
From $295.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the role: As a Principal Security Engineer on the Detection and Response (D&R) team at Roblox, you'll play a key role designing and developing effective custom security data pipeline systems, detection strategies and automations for response workflows to defend our critical assets from threat actors. You will also lead real-time incident response, actively investigate events and analyze threat actor techniques to prioritize emerging threats to ensure Roblox is equipped to mitigate and react to critical challenges. You will play a vital part to ensure the safety of our community and enterprise by proactively fostering a high-performing, inclusive security culture. This is a hybrid in-office role. You Will: Be a D&R authority! You will deliver robust detection & response capabilities: build new threat detection systems (keeping false positives low) while also automating processes with scripts, playbooks and orchestration tooling. Implement ETL pipelines : Design and develop customized data processing pipelines. Conduct security operations : Actively monitor security events and participate in on-call rotations to lead real-time incident response to contain and mitigate potent
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacart’s Enterprise Platform team is responsible for the white-label product that enables retailers of all sizes to compete online through their own ecommerce experiences. The team is focused on platform growth and long-term retailer success — from onboarding new partners to helping them acquire, serve, and retain their end customers in ways that align with their unique brand and strategy. We are seeking a Senior Product Manager to own two important efforts: creative tools that give retailers control over their e-commerce theme and brand experience, accelerated with AI acting as the product lead for large enterprise retailers – shaping proposals, aligning complex retailer requirements, and guiding planning cycles that unlock mutual growth This role is one of the most dynamic and high leverage roles we have. It requires someone who can move fluidly
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Creator Content Platform Team provides a secure, scalable, and extensible foundation for ingestion, processing, storing, managing, and serving user content. As a Principal Software Engineer (Backend, Distributed Systems), you will design and build backend services to help game creators reach the broadest audience possible. You will help us build & scale large distributed systems, data processing and analysis pipelines, an access control system, Open Cloud APIs all of which are central for all content created in Roblox. \ You Will: Solve on a variety of unique technical challenges Have the independence, opportunity and the end-to-end responsibility to design, build, test and deploy services within the Roblox ecosystem Guide the future technical direction of the team and have impact on engineering Be a technical bar-raiser for high code quality, architectural designs, and long-term approaches Mentor and develop fellow engineers on the team Design systems and services that are scalable and resilient Collaborate with passionate, Engineers, Product Managers and other Roblox team members, cross-functionally You have: Experience: You have 13+ years of experience working on backend, d
From $278.5K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. ML Platform @ Roblox today supports hundreds of ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As an Infrastructure Engineer on the ML Platform team, you will design, scale, and maintain the foundational infrastructure powering our entire machine learning ecosystem. We are looking for accomplished engineers to spearhead the development of our next-generation ML tooling and platform capabilities. You will: Bootstrap and maintain Kubernetes and Cloud infrastructure for ML Platform components--Serving Layer, Metadata Store, Model Registry, and Pipeline Orchestrator. Set technical strategy and oversee development of high scale and reliable infrastructure systems. Propose and implement new platform tooling to improve time to production for MLEs and Data Scientists across the full ML lifecycle. Work on infrastructure projects such as GPU fleet management, hybrid-cloud orchestration, and writing custom Kubernetes controllers and resources. Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices. Partner across organizations to build tooling, interfaces, and visualizati
Job Details: Job Description: Join an enthusiastic team of engineers in Intel's Networking Solutions Group (NSG) focused on enabling next generation of programmable Infrastructure Processing Units (IPUs) with our lead customers as part of the Customer Experience Support (CES) organization. Intel brings decades of leadership in networking, virtualization, packet processing, storage, and security to a new class of IPU products that accelerate host networking functions and support emerging use cases such as security, virtualization, storage, load balancing, and data path optimization. Working closely with major cloud service providers and Intel development teams, you will help deliver customized IPU based solutions that enhance isolation, security, performance, storage and system management for our customers. A big part of the day-to-day job is to help customers manage feature request processes, enable solutions, and debug issues. Projects and responsibilities include but are not limited to: • Gain our customers' trust, understand their needs, and build POCs to meet them. Work closely with internal and external partners to understand use cases and requirements. • Be the go-to technical resource for customers building complex Datacenters, AI infrastructure as well as helping them understand performance characteristics for solutions. • Prepare and deliver technical content to customers including presentations, workshops, etc. • Contribute across the full IPU lifecycle, including board and platform bring up, low-level device initialization, OS driver and kernel configuration, system management, feature enablement, use case testing, debugging, and verification. • Defines systems implementation and integration solutions and plans to ensure optimum performance and reliability across hardware, firmware and software w
About the Team OpenAI’s Client Platform Engineering (CPE) team delivers trusted devices at scale: secure by default, reliable by design, and effortless to use. We own platform capabilities across macOS, Windows, iOS, Android, and Linux, spanning endpoint posture and device trust, application delivery, onboarding, updates, telemetry, workflow orchestration, and employee-facing remediation. The team partners deeply with Security, Research, Applied, and specialized engineering groups to enable and protect OpenAI while reducing friction for the people advancing our mission. About the Role As an Engineering Manager for CPE, you will lead a team of engineers responsible for the strategy, delivery, and operation of OpenAI’s cross-platform client foundation. You will combine people leadership with strong technical judgment: setting direction, developing engineers, reviewing architecture and tradeoffs, and creating the operating mechanisms that turn ambiguous needs into durable platform outcomes. This is a high-leverage role at the intersection of security, reliability, developer velocity, and employee experience. CPE is a highly technical platform engineering organization delivering first-party services, automation, observability, and safe fleet operations. We’re looking for a leader who can guide its next chapter, scaling the team and its systems, partnering across the company, and raising the bar for secure, reliable, low-friction experiences across every supported platform. In this role, you will: Lead and develop a high-performing engineering team; hire thoughtfully, coach engineers, create clarity, and foster an inclusive, high-accountability culture that pushes perceived limits. Define and execute a multi-year client-platform strategy and roadmap across macOS, Windows, iOS, Android, and Linux, including how Codex and agents can reshape employee computing. Provide technical direction for endpoint posture, device trust, application delivery, device onboarding, updates,
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Software Engineers build and own high-value products for our customers and the infrastructure that lets our business scale. Vanta's team and technology surface are growing quickly, and it's essential that we invest in the right abstractions and systems to scale with our business. As a Software Engineer, you'll build and own full-stack features and platform primitives across our stack, work closely with the teams and partners who depend on what you ship, and grow into broader technical ownership. Your work will directly accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Integrations Platform team within Vanta Government Cloud mission is to power the world's largest trust automation ecosystem, enabling any person or agent to build, connect, and automate trust seamlessly. Federal compliance is going through its biggest shift in a decade and this team builds the integration platform underneath Vanta Government Cloud that turns federal frameworks into automated, continuously monitored product experiences. This is a builder role: you'll turn federal control requirements into the platform primitives and integrations that make continuous authorization real. We own Vanta's integration ecosystem within the US Government Cloud, which currently includes a growing set of integrations
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Software Engineers build and own high-value products for our customers and the infrastructure that lets our business scale. Vanta's team and technology surface are growing quickly, and it's essential that we invest in the right abstractions and systems to scale with our business. As a Software Engineer, you'll build and own full-stack features and platform primitives across our stack, work closely with the teams and partners who depend on what you ship, and grow into broader technical ownership. Your work will directly accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Integrations Platform team mission is to power the world's largest trust automation ecosystem, enabling any person or agent to build, connect, and automate trust seamlessly. We own Vanta's integration ecosystem, which currently includes over 400 integrations across Cloud Providers (AWS, Azure, GCP), Identity Providers, Mobile Device Management (MDM), and Human Resources Information System (HRIS). We are focused on developing the Integration Platform. This includes creating shared primitives for authentication, lifecycle, observability, and publishing to ensure all integrations are built on the same foundation. Our North Star is to eliminate the barrier to building integrations entirely. We aim to enable any
From $380K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Director, Corporate and Securities Counsel We are seeking a seasoned Corporate and Securities attorney to join our expanding legal team. Reporting to the VP, Deputy General Counsel and Chief Compliance Officer, you will lead securities law and governance matters while managing a small team of legal professionals. You will drive process improvements and operational efficiencies while serving as a strategic business partner to the finance, equity, people, and leadership teams. This role will be based at our San Mateo, CA headquarters (hybrid with Tues-Thurs in office days) and reports into the Deputy General Counsel-Chief Compliance Officer. You will: Direct the preparation and filing of SEC reports (10-K, 10-Q, 8-K) and ensure NYSE compliance, leveraging AI tools to enhance disclosure drafting and consistency. Oversee the preparation and review of earnings releases and public announcements in collaboration with Finance, Investor Relations, and Marketing. Manage board and committee logistics, including agendas, materials, minutes, action items and corporate governance matters. Oversee the management of domestic and international subsidiaries, including intercompany funding, incorporatio
From $399.4K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Content Platform team at Roblox powers the infrastructure behind every asset used across the Roblox ecosystem—enabling creators and developers to bring their visions to life at global scale. From 3D models and images to videos and audio, our platform manages the complete lifecycle of all assets essential for immersive experiences, supporting one of the largest services in the world at over 100+ million requests per second. Our mission is to deliver a seamless, reliable, and innovative content system that empowers creators, supports record-breaking games, and ensures the highest standards of performance and safety for our community. As the Technical Director for Content Platform, you will lead multidisciplinary engineering teams responsible for the technical and product vision of Roblox’s asset infrastructure. You will own the lifecycle of every asset—from creation and upload to storage, indexing, delivery, and rendering in the game client. Your leadership will be critical in scaling our systems, optimizing distributed infrastructure, and enabling new possibilities for creators and players alike. You Will: Define and drive the long-term strategy, architecture, and priorities for the Cont
Other cities to consider
More places hiring for this role
Get new ai platform and agentic engineer jobs in United States by email
Daily job updates · Unsubscribe anytime