NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent engineering manager to own and deliver an end to end manageability stack for Data Center Systems. We are seeking an experienced manager who is deeply technical, hands-on, and has a wide system view. You will manage a team of experts, design & build OpenBMC based manageability software stack for NVIDIA’s next generation Data Center Compute Systems. We want to grow our teams with the smartest people in the world. If you're creative and autonomous, we want to hear from you! What you’ll be doing: Own and deliver OpenBMC based manageability stack for next generation Data Center Compute Systems. Own firmware delivered to data centers in terms of quality, reliability and telemetry performance. Manage and lead a distributed team of software engineers to deliver firmware stack with high quality. Work with data center architects and cloud customers for correct requirements and scope implementation to ensure speed of light product development. Work closely with cross functional teams to ensure scalable manageability architecture for all data centers products Drive efficiency, reliability and optimization in firmware architecture from a data center view point. Work closely with customers and internal teams to resolve issues at Speed of Light. What we need to see: BS, MS, or PhD in EE/CS or related field o
Jobs in United States
Software Engineer Ml Infrastructure Platform in United States
2,123 active opportunities · Updated October 2026
Showing
15 jobs
Explore current software engineer ml infrastructure platform jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About Pinecone: Pinecone is the leading vector database for building accurate and performant AI applications at scale in production. Pinecone’s mission is to make AI knowledgeable. More than 109,000 customers across various industries have shipped AI applications faster and more confidently with Pinecone’s developer-friendly technology. Pinecone is based in New York and has raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Role/Team: As a Technical PMM , you will help developers, software engineers, and machine learning scientists understand how Pinecone enables knowledgeable AI applications. You will accomplish this through messaging and positioning, GTM strategy, launch planning, and enablement for key product features. You will work closely with Product, Sales, Developer Relations, and Marketing teams to ensure our customers understand the value of our platform. Responsibilities: Product Launches & GTM Strategy: Develop launch plans for product releases, align with product teams, and manage GTM execution. Lead core launch teams with cross-functional members (documentation , communications, developer relations, product, growth marketing, etc). Measure customer impact and adoption post-launch, iterating on strategies as needed. Positioning & Messaging: Craft compelling, technically accurate messaging for AI-powered search and retrieval solutions. Build product narratives for enterprise and developer audiences that differentiate Pinecone from competitors. Develop and maintain positioning for new product capabilities Customer & Market Insights: Partner with customers to develop case studies, gathering business impact metrics and architectural insights. Engage in customer interviews and research to refine messaging and identify key value drivers. Content Development & Enablement: Develop internal and external enablement materials, including sales training, pitch decks, and GTM enablement sessi
About the Team At OpenAI, we are dedicated to building safe artificial general intelligence (AGI) to benefit all of humanity. Our mission attracts the world’s top talent in science, engineering, and business to address one of the most ambitious challenges of our times. The Recruiting team is at the heart of this mission, tasked with identifying and hiring exceptional individuals who align with OpenAI's values and cultural ambitions. Our approach to recruitment aims to set the standard for excellence and innovation in the field, connecting outstanding candidates with opportunities to impact the future of AI. About the Role As a Senior Technical Sourcer at OpenAI, you will play a key role in identifying and engaging top-tier engineering talent across a broad range of technical domains. Your focus will be on sourcing exceptional software engineers and technical talent to help build world-class teams advancing our mission in AI research and deployment. In this role, you will: Lead sourcing strategies to identify and engage candidates across a variety of engineering disciplines, including software engineering, backend systems, product engineering, and related technical areas. Develop and maintain a strong pipeline of passive candidates through proactive outreach, research, and networking. Collaborate closely with hiring managers and technical leaders to deeply understand hiring needs and refine sourcing approaches accordingly. Leverage advanced sourcing techniques and tools to identify and attract exceptional talent. Represent OpenAI at industry events and conferences to promote our mission and connect with potential candidates. Maintain accurate and organized candidate data and metrics to inform sourcing strategies and decision-making. You might thrive in this role if you have: 5+ years of experience in technical sourcing, with a focus on engineering or technical roles. A proven track record of successfully sourcing candidates across a range of software engineering and
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. The Grid Service and Platform Engineering team is looking for a highly motivated and collaborative Software Engineering Manager. This role involves leading the engineering of mission-critical, tier 0 service infrastructure, the foundational data platform that powers Smartsheet at scale. You will oversee services that handle millions requests per day, operate at 99.999% availability, and deliver low-latency, high-throughput performance for millions of customers worldwide. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to the Director, Engineering and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Manage one or more related teams of 6–10+ software engineers, driving development of tier 0 grid services and platform infrastructure that millions of customers depend on daily. Own and uphold 99.999% service availability targets across critical platform services, embedding reliability engineering, incident management, and on-call rigor into team culture. Help architect and guide technical vision to evolve low-latency, high-throughput service platforms capable of sustaining millions requests per day with predictable, consistent performance under load. Guide and mentor engineers on distributed systems architecture, scalability patterns, and platform best pr
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Executive Director, Technology Product Management – Medical Cost Initiatives Location: Remote Department: Healthcare Technology & Product Management Reports To: VP, Chief Health Informatics Officer (CHIO) - CVS Healthcare Delivery Role Overview As the Executive Director, Digital Product – Medical Cost Initiatives , you will lead the technology product strategy and execution for key medical cost portfolios for CVS Healthcare Delivery businesses. Core areas of focus are Care Transitions, Patient Segmentation, Chronic Condition Pathways, and Specialty Care Services. In this high-impact executive role, you will partner directly with clinical and business operational leaders to shape end-to-end technology pathways that support CVS Healthcare Delivery care models. You will lead a multidisciplinary team of product managers, Epic analysts, software engineers, data scientists, and UI/UX designers to translate clinical vision into scalable technology roadmaps that optimize patient outcomes and lower the total cost of care. The ideal candidate must have expertise with value-
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Engineering Manager, Cloud Efficiency Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. We are looking for an experienced Engineering Manager to lead the Cloud Efficiency engineering team. In this role, you will own the technical vision and execution for building a unified, self-serve cloud efficiency platform along with AI skills and agents that makes resource usage and spend attributable and governable while driving insights and optimization of our cloud spend. AS AN ENGINEERING MANAGER IN CLOUD EFFICIENCY, YOU WILL: Lead and grow our talented team of software engineers, fostering a culture of technical excellence, ownership, and continuous learning. Drive the roadmap for Cloud Efficiency — translating company-level spend objectives into engineering systems: authoritative cost data, resource ownership registry, attribution pipelines, cost and unit economics modeling, observability, governance policies, and optimization workflows — in partnership with Product, Engineering, Finance, and Data Science. Set technical strategy for backend systems, data pipelines, and APIs that measure, attribute and surface cost and usage insights at
From $198K/yr
Chicago Trading Company (CTC) is a premier proprietary trading firm specializing in options market making. Our collaborative culture fuels innovation in quantitative research, systematic trading strategies, and cutting-edge trading technology. For over three decades CTC has provided critical liquidity across derivatives exchanges worldwide - making them fairer, more transparent, and more efficient. We strive to be the most innovative firm in the industry today, tomorrow, and long into the future while upholding ethical excellence. We believe that CTC makes a positive impact on the markets, the lives of our employees, and all the communities to which we belong. Started in 1995 by a team of forward-thinking Traders, we are proud to call ourselves an industry leader that keeps making markets and each other better. The Role As a Quant Trading (QT) Intern, you will be challenged to learn and adapt in an exciting team environment and will play a substantial role in our day-to-day trading and quant-related activities. Your impact is immediate and meaningful. You will become a member of a team for the summer and play a vital role completing project work and/or identifying trading opportunities and communicating with software engineers, quants, traders, and risk managers throughout the trading day. As part of the Summer Associate cohort, you will learn and socialize alongside other QT Interns and Software Engineering (SE) Interns. What to Expect The 8-week internship program gives you insight into our culture and an inward look into what our business is all about. You will participate in three classroom learning experiences including: a week-long Basics of Options class, a multi-week class on the fundamentals of market making (Mock Trading), and a multi-week class that introduces elements of our quant framework (Quant Curriculum). You will also attend planned social activities and talks from various business leaders to propel your professional growth, and
About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse, misalignment, and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. About the Role We are hiring a Product Manager to focus on risk related to multimodal models. In this role, you will drive initiatives which ensure that OpenAI’s audio, image, and video deployments are safe, impactful, and aligned with user needs and technical innovation. You will clarify strategic priorities, develop safety-focused product roadmaps, and collaborate closely with AI researchers, software engineers, policy experts, and cross-functional partners. This role suits a proactive, technically skilled product manager adept at adversarial thinking and excited to tackle challenging, ambiguous problems through structured analysis and collaborative decision-making. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Partner closely with AI research, engineering, data science, policy teams, and other stakeholders to embed safety throughout the development and deployment of multimodal AI models - such as GPT-Live and ChatGPT Images - as well as multimodal capabilities in frontier AI models. Develop comprehensive frameworks for understanding and mitigating deployment safety risks, drawing on data analysis, expert consultation, and adversarial assessments. Define strate
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Location: Bellevue, WA Engineering Manager, Cost Intelligence We are looking for an experienced Engineering Manager to lead the Cost Intelligence engineering team. In this role, you will own the technical vision and execution for the features and systems that help Snowflake customers understand, monitor, and optimize their Snowflake consumption and spend. You will lead a talented team of engineers building the data products, APIs, and platform services that power cost visibility, usage analytics, budgeting, and cost optimization insights across Snowflake's platform. You'll work closely with Product Management, Design, Data Science, and cross-functional engineering teams to ship world-class cost intelligence capabilities to Snowflake's customer base. As manager for the Cost Intelligence team, you will: Lead and grow our talented team of software engineers, fostering a culture of technical excellence, ownership, and continuous learning. Drive the roadmap for Cost Intelligence features — including cost allocation, resource budgeting, anomaly detection, and optimization recommendations — in partnership with product management. Set technical strategy for backend systems, data pipelines, and APIs that surface cost and usage insights to customers at massive scale. Own delivery end
From $280.5K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Product Manager in Engineering Acceleration , you will define the vision and strategy for how software engineers work at Roblox. Your mission is to ensure that Roblox engineers spend more time building the metaverse and less time managing infrastructure complexity in your areas of responsibility. The Engineering Acceleration portfolio includes Source Control, Testing, Secure Software Supply Chain Management, Continuous Deployment, and Observability among many others. In this role, you will primarily own the developer experience for Continuous Deployment, Testing , and Observability , while remaining adaptable as organizational priorities evolve. AI has been transforming how we approach these systems. We’re already leveraging AI to test our software and to identify and diagnose production incidents. The successful candidate here will bring deep expertise not only in the software development lifecycle, but crucially also on the rapidly evolving landscape of AI tooling. This is a rare opportunity for an infrastructure product leader to drive meaningful impact at scale across a very large engineering organization. You Will: Define and drive the long-term visi
From $192K/yr
As Engineering Manager for Threat Detection, you will lead a high-performing team that powers Datadog's detection program. Threat Detection is the organization responsible for keeping Datadog ahead of an evolving threat environment: closing coverage gaps faster, raising the bar on signal quality, and shipping detections that hold up under the scale and complexity of cloud-native infrastructure. Your team will combine direct detection expertise, platform engineering, and applied AI to ship detections at a pace and scale traditional rule-writing alone cannot match. Examples of what your team will work on include detection-authoring agents, the detection platform that powers every rule in production, coverage analysis, alert triage and response automation, and the evaluation infrastructure that holds these systems to a high bar of fidelity. Detection authorship is a shared responsibility across the organization, and your team will contribute both by building the systems that scale our authoring capacity and by writing detections directly when their domain expertise is the right tool. You will partner closely with our Security Incident & Response Team (SIRT), Cyber Threat Intelligence (CTI), AI Engineering teams, and Datadog's broader Security organization. This is a high-impact leadership role: you will grow a team of security and software engineers responsible for building and executing our detection and AI strategy. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the strategy, roadmap, and execution of Datadog Security's shift to AI-accelerated detection and response. Drive development of high-fidelity detections as a shared responsibility across the organization, ensuring your team's systems and direct contributions raise the bar on coverage and
Software Certification Engineers (Senior, Lead, or Expert) (Virtual) — United States - Remote. Apply via Workday.
As one of the technology industry's most desirable employers, NVIDIA has been redefining accelerated computing, computer graphics and leading the Artificial Intelligence revolution. NVIDIA's innovation is fueled by its great technology—and amazing people. We are seeking a Senior Silicon and System Product Lead to influence, innovate and take our next generation products to the market. As part of the Silicon Solutions Team, we architect and deliver groundbreaking system solutions that integrate all aspects of the system from silicon design, software design to operations and final deployment in multiple market segments that NVIDIA serves. This position offers an unique opportunity to collaborate with multiple organizations in the company and grow your career in a high impact role. We need a passionate, hard-working and creative individual to lead the products all the way from market analysis to delivering the features on the final product. What you'll be doing: Drive product performance and power targets, trade-off features/configurations and provide innovative solutions to complex silicon and system level problems. Evaluate new market segments and use cases; translate market requirements to engineering problem statements and metrics. Innovate Performance, power, yield and quality optimizations and features for the world’s fastest power-shipping products in the GPU and SoC market segments spanning gaming, automotive, datacenter and DL/AI. Develop methodologies and requirements for multi-functional teams to drive silicon and system product features to production. Incorporate productization feedback to improve the next generation. Lead the team for feature requirements and schedule from architecture to silicon phase of projects. Work alongside system architects, designers, marketing teams, chip and board designers, software/firmware engineers, HW/S
Our work at NVIDIA is dedicated towards a computing model focused on visual and AI computing. For two decades, NVIDIA has pioneered visual computing, the art and science of computer graphics, with our invention of the GPU. The GPU has also shown to be spectacularly effective at solving some of the most complex problems in computer science. Today, NVIDIA’s GPU simulates human intelligence, running deep learning algorithms and acting as the brain of computers, robots and self-driving cars that can perceive and understand the world. We are looking to grow our company and teams with the smartest people in the world and there has never been a more exciting time to join our team! The AI Infrastructure Product Design team creates software used by engineers and researchers to prepare data, run complex workflows, and understand results. This internship offers ownership of a defined product problem from early research through a tested design and implementation handoff. Designers on this team often move between Figma and working HTML prototypes, and may hand off HTML directly to engineering. This work calls for a high standard of visual and interaction design alongside technical fluency. The role is a good fit for someone who enjoys making technically complex systems easier to understand and who uses large language models and software agents thoughtfully as part of their design and prototyping process. What you will be doing: Own a focused design project for an internal AI infrastructure product, from understanding the problem through a validated design and implementation handoff. Interview engineers and researchers, map their workflows, and turn the findings into clear product requirements, user flows, and interaction models. Create precise, implementation-ready interface designs and interactive prototypes in Figma and HTML/CSS, with careful attention to typography, hierarchy, spacing, visual consistency, interactio
From $110K/yr
Datadog AI Research — Scholars Program with Carnegie Mellon University Datadog AI Research (DAIR) is partnering with Carnegie Mellon University to support a small number of PhD students working on open research problems grounded by ongoing efforts at Datadog/DAIR. You will frame a problem, run your own experiments, and write up what you find, with compute and data at a scale most academic labs cannot provide. You will collaborate with colleagues working on the same questions. The Lab And The Research: DAIR is an industrial research lab motivated by practical challenges in observability and software operation: detecting and diagnosing failures, understanding complex production environments, and helping engineers operate software more effectively. The lab focuses on creating specialized foundation models, post-training and evaluating AI agents, and building frontier-scale machine learning systems. By combining fundamental research with Datadog's large-scale, real-world data and infrastructure, the lab develops new AI capabilities and translates them into practical systems with meaningful impact. Internship projects are shaped with your DAIR mentor and your CMU faculty advisor. You do not need prior experience with observability, monitoring, or infrastructure. What You'll Do: Own a research project end to end: framing the question, running the experiments, writing it up Work directly with a DAIR mentor engaged in the same problem, and stay connected to your advisor and lab Publish, and use the work toward your dissertation See research reach production, when it works Who You Are: Currently enrolled in a PhD program at Carnegie Mellon in machine learning, computer science, statistics, or a related field Depth in at least one area relevant to the research above Comfort running real experiments — training models, working with GPUs, reading and reimplementing recent papers Evidence you can do research: conference or workshop papers, preprin
Higher-paying openings
Jobs with higher listed pay
Staff Software Engineer - Fern
Postman · New York, California, United States
Staff Software Engineer, Business Platform
Postman · San Francisco, California, United States
Principal Software Engineer
Roblox · San Mateo, CA, United States
Principal Software Engineer, Game Safety
Roblox · San Mateo, CA, United States
Staff Software Engineer- Codegen
Postman · Austin, Texas, United States
Sr. Staff Software Engineer, Merchants
Pinterest · San Francisco, CA, US
Related career options
Similar roles with stronger pay
Demand 46/100 · 8 jobs
$840K – $840K/yr
Salary →Demand 43/100 · 6 jobs
$382.5K – $382.5K/yr
Salary →Demand 43/100 · 8 jobs
$300K – $300K/yr
Salary →Demand 42/100 · 7 jobs
$300K – $300K/yr
Salary →Demand 38/100 · 30 jobs
$278.9K – $278.9K/yr
Salary →Demand 30/100 · 11 jobs
$255.7K – $255.7K/yr
Salary →Other cities to consider
More places hiring for this role
Get new software engineer ml infrastructure platform jobs in United States by email
Daily job updates · Unsubscribe anytime