NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are looking for an outstanding Compiler Engineer to help build the next generation of intelligent compiler technologies for NVIDIA's accelerated computing stack. Our team works at the intersection of compilers, agentic systems, numerical correctness, and verification to create systems that can reason about, generate, optimize, and validate code transformations across software and hardware boundaries. This is an excellent opportunity for new graduates who are excited about coding agents, AI-assisted software engineering, developer tools, and GPU computing. In this role, you will work with experienced engineers and researchers to build agentic systems and compiler-aware tooling that improve developer productivity, code quality, and system performance across NVIDIA's software and hardware stack. What you'll be doing: Build and improve coding-agent systems for tasks such as code generation, transformation, debugging, optimization, validation, and developer assistance. Develop agent workflows involving tool use, planning, memory, execution, and feedback loops for software engineering and compiler-related tasks. Help create training, evaluation, and verification environments to improve agent quality, correctness, r
Jobiba hiring network
Ai Systems Engineer Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current ai systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. The opportunity The Autonomous Freight Systems team is a brand new, AI-first engineering team in San Francisco. This team will own Flexport’s client-facing rates platform and the self-serve freight booking experience—two of the highest-leverage surfaces in our Client App that dictate how clients see pricing and book freight without manual intervention. As a Staff Engineer, you will be the technical anchor for our next-generation AI-powered rates platform. We aren't just building a UI; we are building AI agents that handle real logistics work: parsing complex rate sheets, managing pricing intelligence across ocean, air, and trucking, and making "Self-Serve" a reality for thousands of shippers. You will build the intelligence layer that allows clients to commit freight on technology alone, with no account executive and no operations touch. This is a ground-floor opportunity to shape technical direction, set up a new codebase, and build the platform that moves Flexport from an assisted-sales model to a tech-run one for the long tail of our client base. You will partner with Pricing, Sales, Ops, and Design on hard domain problems and ship to the most-used client
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Luau App Foundations team is responsible for the core infrastructure of the Roblox application. They bridge the gap between app and the performance-heavy Roblox game engine. This team is building the unified stack that powers mission-critical surfaces like Home, Avatar, and Social for millions of concurrent users. They are the ones who make it possible for "Web-style" development efficiency to exist within a high-performance C++ game engine. Why is this role exciting: Technical Pioneer: You will be writing libraries and modules using C++ inside a world class Roblox proprietary Game Engine. Systematic Impact: This is a "Force Multiplier" role. The frameworks and components you build will be used by dozens of other engineering teams to ship their features. Complex Problem Solving: You aren’t just building an app; you’re managing smooth data flow through a client that has to perform perfectly on a $100 Android phone and a $3,000 Gaming PC simultaneously. 0 to 1 Transitions: You will lead the charge in shaping some of the most crucial components of the App written in C++ inside the Game Engine. Key Challenges: Bridging Tech Stacks: The libraries you write sit between a modern UI (written in
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Corporate Systems team focuses on maintaining secure, reliable systems that support employees across the company. This team works closely with Security, IT, and Engineering partners to manage identity systems, endpoint devices, and cloud infrastructure. The goal is to ensure systems are scalable, dependable, and prepared to address evolving security risks. As an Application Engineer, you will manage and improve systems that support identity, device management, and cloud infrastructure. You will build automation to streamline account and device lifecycle processes, respond to technical issues, and contribute to engineering standards through code reviews and documentation. You will also use AI-assisted tools to enhance development workflows and improve efficiency. This role is based in our Menlo Park, CA office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do You manage endpoint devices, identity systems, Okta, Go
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Software Engineer, Distributed Data Systems, you will design, build, and operate some of the largest distributed data systems in the world. You will be responsible for the end-to-end stack to deliver and consume top-quality data for robotics training at exabyte-scale. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI’s rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable, large-scale systems in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure such as exabyte-scale distributed data processing, data selection, automated labeling, and training data loaders. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. Deliver the best possible data for training robotics models. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented a
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI's rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable infrastructure in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented and bring rigor to building and maintaining reliable systems. Demonstrate excellent software enginee
About the Team The Intelligence & Investigations Engineering team builds systems that detect, analyze, and disrupt abuse across OpenAI’s products. We partner closely with the Child Safety team and cross-functional groups to protect users while advancing OpenAI’s goal of developing AI that benefits everyone. About the Role As a Fullstack Engineer focused on child safety, you’ll build data-intensive, AI-powered applications and infrastructure that enable operators and investigators to work effectively and responsibly. You’ll adapt quickly in ambiguous, fast-moving environments to deliver well-crafted, reliable tooling for high-severity safety work. *Candidates should understand this role involves exposure to sensitive and egregious content. In this role, you will: 4+ years of experience as a software engineer Prototype, build, and maintain intelligence systems that detect, triage, and enable efficient human review of possible high severity harm Work hand in hand with operators and investigators, designing and delivering systems that enable them to do their work faster, more accurately, and more safely. Develop across the stack: UIs, services, pipelines, and anything else required to solve the problems we face. Interact with partners across Product Policy, Platform Integrity, Safety Systems, and Research Contribute to the team’s technical strategy, especially for child safety related tools and systems Report on impact in a data-driven fashion You might thrive in this role if you: Have a strong software engineering foundation and enjoy owning systems end-to-end—from infrastructure and data ingestion to frontend tooling Are energized by working at the frontier of AI capabilities, integrating new models and APIs into practical systems Have experience building and operating large-scale data pipelines or search/retrieval systems Are proficient in Python and/or TypeScript, and familiar with tools like Spark, Kafka, Flink, data warehouses, and SQL Take a product-minded ap
Location: San Francisco, CA (Remote/Hybrid Available) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of the brightest industry experts in the field building cloud-native applications that scale to trillions of data points collected from electricity markets globally. You will be a part of a dynamic, robust team primarily supporting the backend needs of our Aria software product spanning hundreds of data sources, sinks, services, and jobs. Your expertise will not only have a direct impact on product decisions, but you also be well-positioned to drive the development and trajectory of our entire platform and infrastructure and influence important architectural decisions that affect the whole organization. Key Responsibilities Foster a culture and mindset of well-designed systems, test-driven software, and transparent communication with a high caliber of mutual respect and consideration for stakeholders Read and write a lot of Go, Python, and Protobuf Build, test, debug, maint
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e
MongoDB’s mission is to empower innovators to create, transform, and disrupt industries by unleashing the power of software and data. We enable organizations of all sizes to easily build, scale, and run modern applications by helping them modernize legacy workloads, embrace innovation, and unleash AI. Our industry-leading developer data platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available in more than 115 regions across AWS, Google Cloud, and Microsoft Azure. Atlas allows customers to build and run applications anywhere—on premises, or across cloud providers. With offices worldwide and over 175,000 new developers signing up to use MongoDB every month, it’s no wonder that leading organizations, like Samsung and Toyota, trust MongoDB to build next-generation, AI-powered applications. Atlas Search is a multi-cloud service that allows users to execute complex full text and vector search queries using the MongoDB Query Language . Our users are free to focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team builds and maintains the instances and supporting infrastructure powering Atlas Search. This platform deploys and monitors search deployments, providing a highly scalable yet observable system for customers and engineers. The Atlas Search product is quickly gaining traction with customers and we are shipping core infrastructure components that enable this growth. This role is based in San Francisco, CA with an in-office or hybrid work model. Successful candidates will have the following qualities: 2+ years of hands-on experience designing, building, testing, and maintaining industrial-strength backend software and automation in complex codebases Experience developing distributed systems and multithreaded applications Familiarity with public cloud platforms, distributed infrastructure, and metric-based development Experience with at least one modern statically typed program
Atlas Search is a multi-cloud service that allows users to execute complex full text and vector search queries using the MongoDB Query Language . Our users are free to focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building the cloud-based distributed systems software responsible for the lifecycle of search indexes including: data ingestion, index building, partitioning, performance, availability, and backup management. Our product is quickly gaining traction with customers and we are making core architectural improvements that you will contribute to. This role is based in Toronto, ON hybrid. Successful candidates will have the following qualities: 2+ years of hands-on experience designing, building, testing, and maintaining industrial-strength backend software in a complex codebase Experience developing distributed systems and cloud services Experience with at least one modern statically typed programming language, and interest in working with Java Excellent verbal and written technical communication skills and enthusiasm for collaborating closely with colleagues A growth mindset and the desire to learn quickly through taking on challenges, reflecting on outcomes, and incorporating feedback A strong sense of ownership over their work, from initial design all the way through maintaining code in production You will: Contribute to the design, implementation, and support of projects that improve the scalability of Atlas Search to make using it a seamless experience for even the largest workloads Work with a collaborative team that prioritizes sound technical decision-making and building systems that our customers love and that we are proud of as engineers Have the opportunity to lead projects and own subsystems Provide input on the team’s roadmap and help determine the architecture of our system Success measures: In 3 months you’ll have a solid high-level understanding of what our team does and how we operate.
Atlas Search is a multi-cloud service that allows users to execute complex full text and vector search queries using the MongoDB Query Language . Our users are free to focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building the cloud-based distributed systems software responsible for the lifecycle of search indexes including: data ingestion, index building, partitioning, performance, availability, and backup management. Our product is quickly gaining traction with customers and we are making core architectural improvements that you will contribute to. We are looking to speak to candidates who are based in San Francisco, CA for our hybrid working model. Successful candidates will have the following qualities: 2+ years of hands-on experience designing, building, testing, and maintaining industrial-strength backend software in a complex codebase Experience developing distributed systems and multithreaded applications Experience with at least one modern statically typed programming language, and interest in working with Java Excellent verbal and written technical communication skills and enthusiasm for collaborating closely with colleagues A growth mindset and the desire to learn quickly through taking on challenges, reflecting on outcomes, and incorporating feedback A strong sense of ownership over their work, from initial design all the way through maintaining code in production You will: Contribute to the design, implementation, and support of projects that improve the scalability of Atlas Search to make using it a seamless experience for even the largest workloads Work with a collaborative team that prioritizes sound technical decision-making and building systems that our customers love and that we are proud of as engineers Have the opportunity to lead projects and own subsystems Provide input on the team’s roadmap and help determine the architecture of our system Success measures: In 3 months you’ll have a
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a principal engineer on the Online Systems team, you’ll join a team that powers Pinterest’s most business-critical online systems at massive scale, driving the reliability, efficiency, and evolution behind every core Pinner and Advertiser experience. You'll lead major efforts like multi-region deployment and Kubernetes migration, set the standard for operational excellence, and define the long-term vision for our online serving infrastructure, supporting machine learning and product innovation across the company. This is an opportunity for high-impact technical leadership, broad visibility, and cross-functional influence at the heart of Pinterest’s platform. What you’ll do: Improve reliability, scalability and infra efficiency for Pinterest’s critical online systems across storage and caching, online service and realtime analytics syste
About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b
Get new ai systems engineer jobs by email
Daily job updates · Unsubscribe anytime