Jobiba hiring network

Ai Systems Engineer Jobs

10,000 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current ai systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto

pythonci/cdgit
View job →

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,

pythonkubernetesmachine learning
View job →
R
Roblox
📍 San Mateo• Full-time• From $293.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Engine Networking Team pulls the players together by ensuring the communication of the game state to all. As a Principal Engineer on this team you will help the players experience the game as a nearly synchronous world. The networking and asset loading team plays a key role in a smooth experience for the players. You will work in all areas of the game platform in your quest for real-time communication of every part of Roblox. You Will: Lead engineers with 8+ years of industry experience Be experienced with one of these area: asset loading, rendering, and networking coming from a Game Engine/Studio. Be an amazing systems-level C++ programmer and be fascinated by the actual work the CPU does when you use smart pointers, templates, virtual functions, and blocks of memory, both structured and raw Have a keen to each millisecond of the network exchanges: You know where the time goes and how to reduce the waste Understand what happens on the operating system level when certain code is completed You Have: Worked on the guts of a multi-player game engine, solving problems related to scale, performance, latency, and throughput in client/server environments. Worked on a very large multithreaded d

awsgitai
View job →
HI
HP IQ
📍 San Francisco• Full-time• C$45 – C$51/hr
19 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Software Engineering Intern on the Systems team at HP IQ, you’ll work on low-level software that sits close to the hardware and helps power intelligent experiences across our products. This role is ideal for students who enjoy understanding how complex systems work under the hood. You’ll have the opportunity to work across multiple layers of the software stack, investigate performance bottlenecks, optimize system behavior, and build software that interacts closely with hardware and system resources. We’re looking for engineers who are curious about more than whether something works — you want to understand how it works, why it performs the way it does, and how to make it better. What You Might Do Build and optimize low-level systems software using languages such as C and C++. Investigate performance bottlenecks and improve the speed, efficiency, and reliability of existing systems. Work on data processing and sensor pipelines that connect software with underlying hardware. Analyze and improve memory usage, resource management, and system performance. Work across multiple lay

redislinuxrest
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role Operating a vehicle remotely over cellular networks is challenging and critical. You will be responsible for ensuring that our "eyes on the road" never blink. You’ll tackle deep-stack networking challenges—from bonding multiple LTE carriers to designing custom FEC (Forward Error Correction) algorithms that out-perform standard protocols. About the Work Engineered Connectivity: Architect a network bonding framework to aggregate bandwidth across multiple cellular providers (Verizon, AT&T, T-Mobile) to ensure zero-drop connectivity. Performance Modeling: Build sophisticated ns-3-like simulations to "stress test" our stack against edge cases like tunnel entries, rural dead zones, and network congestion. Optimization: Develop and implement custom congestion control algorithms specifically tuned for high-bitrate, low-latency video streaming. Cross-Functional Leadership: Partner with Hardware and Embedded teams to optimize the netw

linuxaic++
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role Operating a vehicle remotely over cellular networks is challenging and critical. You will be responsible for ensuring that our "eyes on the road" never blink. You’ll tackle deep-stack networking challenges—from bonding multiple LTE carriers to designing custom FEC (Forward Error Correction) algorithms that out-perform standard protocols. About the Work Engineered Connectivity: Architect a network bonding framework to aggregate bandwidth across multiple cellular providers (Verizon, AT&T, T-Mobile) to ensure zero-drop connectivity. Performance Modeling: Build sophisticated ns-3-like simulations to "stress test" our stack against edge cases like tunnel entries, rural dead zones, and network congestion. Optimization: Develop and implement custom congestion control algorithms specifically tuned for high-bitrate, low-latency video streaming. Cross-Functional Leadership: Partner with Hardware and Embedded teams to optimize th

linuxaic++
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity The Telemetry Data Platform group at New Relic builds the foundation for all of our products: data ingest, storage, and query. As an engineer working on NRDB, you’ll be contributing directly to the proprietary telemetry database technology at the core of our business. We own our software from top to bottom and are directly responsible for its quality and reliability. Each member of the team shares our pager rotation and will occasionally be on-call to respond to system failures; so we prioritize work that keeps the lights on and the pager quiet, in addition to the work that powers all of our new products and streams of data. If the idea of working on systems that process millions of messages per second and handle exabytes of data excites you, then you may be an excellent fit! What you'll do Develop new features with a focus on optimizing performance and efficiency Collaborate with the team to implement scalable solutions and enhance application performance Identifying and acting on opportunities to improve the reliability of our services This role requires 2+ years of professional experience in distributed SaaS software development. Proficiency in Java programming, expertise with algorithms and data structures, and building high-throughput software following best-practices. Deeper understanding of distributed systems and their core challenges. Experience using the command line to manage, investigate, and fix things when they’re broken. Expe

javasqlmysql
View job →
R
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on Creator Services Data, you’ll be leading the company’s efforts to build the next generation Data Storage systems to power the millions of experiences on the Roblox Platform. We run the mission critical cloud services, Data Stores , Memory Stores , and Badges , which are crucial for storing game state such as inventory and scores, implementing leaderboards, server lists, and trading, and tracking player progress and achievements. Our team is also responsible for building dashboards to provide insights to Creators using cloud services including Client/Server Performance , Data Stores , and Memory Stores . Finally, our team owns the Roblox Extended Services platform, which provides the capability for large experiences to purchase additional resources for existing services like Data Stores and new services built around compute and generative AI. At its core, this team is focused on solving complex back end distributed systems and storage problems at scale. However, our scope extends to full stack projects spanning all the way from the infrastructure layer, through data storage and data pipelines, microservices, telemetry, game servers,

awsazuregcp
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join our Fusion product team and help redefine how enterprises use storage. Pure Fusion is an industry-first autonomous storage delivery platform that brings cloud-like agility to enterprise environments by unifying fleets of Pure arrays into a consistent, policy-driven experience. Through a single, integrated control plane, customers can provision and manage storage across their environment - reducing manual complexity, improving consistency, and freeing teams to focus on innovation. As part of the team, you’ll build the distributed systems, APIs, automation, and intelligent services that make this experience possible. You’ll help turn storage into an on-demand, self-service resource that is easier to deploy, scale, and manage across FlashArray, FlashBlade, and Pure Storage Cloud. This is an opportunity to solve challenging problems at the intersection of cloud infrastructure, data, and automation. And to shape a platform that is changing how modern businesses manage data. WHAT YOU'LL DO Develop resilient distributed control systems that support reliable service delivery at scale. Create and maintain API libraries that enable consistent integration across the platform. Design and build declarative, intent-based policy management engines. Implement scalable transactional processing for complex, distributed environments. Integrate platform capabilities with products such as FlashArray, FlashBlade, and Pure1. Collab

pythonjavaangular
View job →
R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the Roblox platform. In the Roblox Engine, the DataModel is a tree-like structure that is analogous to a scenegraph in other 3D engines. This role will report to the engineering manager and will be based out of our HQ in San Mateo, CA in a hybrid model 3 days a week (Tuesdays to Thursdays). Our team owns: The core structures and systems are used to build the DataModel and interact with it. The C++ reflection bindings that form the Engine’s Luau API surface and let creators interact with the DataModel. We’ve built custom codegen tooling to generate the C++ for these reflection bindings and other related structures. DataModel serialization … and much more! You will: Develop engine code that performs well for all user-created games on the Roblox platform. Build the core systems and data structures used in the Roblox engine, working with other teams to find universal solutions. Take ownership of projects throughout their full lifecycles. Execute code that performs well on all the devices Roblox supports—from desktop clients to mobile phone clients to

awsgitai
View job →
R
Roblox
📍 San Mateo• Full-time• From $242.1K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the Roblox platform. In the Roblox Engine, the DataModel is a tree-like structure that is analogous to a scenegraph in other 3D engines. This role will report to the engineering manager and will be based out of our HQ in San Mateo, CA in a hybrid model 3 days a week (Tuesdays to Thursdays). Our team owns: The core structures and systems are used to build the DataModel and interact with it. The C++ reflection bindings that form the Engine’s Luau API surface and let creators interact with the DataModel. We’ve built custom codegen tooling to generate the C++ for these reflection bindings and other related structures. DataModel serialization … and much more! You will: Develop engine code that performs well for all user-created games on the Roblox platform. Build the core systems and data structures used in the Roblox engine, working with other teams to find universal solutions. Take ownership of projects throughout their full lifecycles. Execute code that performs well on all the devices Roblox supports—from desktop clients to mobile phone clients to con

awsgitai
View job →
R
Roblox
📍 San Mateo• Full-time• From $242.1K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the Roblox platform. In the Roblox Engine, the DataModel is a tree-like structure that is analogous to a scenegraph in other 3D engines. This role will report to the engineering manager and will be based out of our HQ in San Mateo, CA in a hybrid model 3 days a week (Tuesdays to Thursdays). Our team owns: The core structures and systems are used to build the DataModel and interact with it. The C++ reflection bindings that form the Engine’s Luau API surface and let creators interact with the DataModel. We’ve built custom codegen tooling to generate the C++ for these reflection bindings and other related structures. DataModel serialization … and much more! You will: Develop engine code that performs well for all user-created games on the Roblox platform. Build the core systems and data structures used in the Roblox engine, working with other teams to find universal solutions. Take ownership of projects throughout their full lifecycles. Execute code that performs well on all the devices Roblox supports—from desktop clients to mobile phone clients to con

awsgitai
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Our mission is to enable effective financial decisions through reliable data, increased efficiency, and automation. We support Marketing, Sales, Seller, Accounting, Tax, Finance and Strategy (F&S), Finance Operations (FinOps), and Treasury functions across automation, data insights, and process improvements. You may work on a wide variety of critical business areas including Seller Systems — Responsible for building the systems and tooling that make sellers and internal revenue teams at Stripe dramatically more productive and effective. We partner with Sales, Finance, Legal, and Product to deliver a single "plane of glass" selling experience that spans deal creation and modeling, negotiation and approvals, contracting, onboarding, and activation. The Seller Systems team composes first‑party, custom Stripe components with best‑in‑class third‑party business systems to deliver configurable, auditable, and globally scalable workflows. Engineers on Seller Systems build services, APIs, integrations, data pipelines, and internal UIs that power seller productivity, reduce time‑to‑activation, improve deal velocity, and enable AI‑driven assistive workflows. Use cases include deal modeling and pricing engines, approval and orchestration platforms, CLM/CPQ integrations, onboarding automation, seller analytics, and AI‑assisted seller tooling. Finance Engineering — Responsible for building the robust and scalable infrastructure that powers

sqlmysqlaws
View job →
O
1mo ago

About the Team This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. About the Role As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. We expect you to: Build and operate reliable infrastructure for research workloads and research-facing services. Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and service reliability layers.

awskubernetesci/cd
View job →

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

pythonsqlaws
View job →
🔔

Get new ai systems engineer jobs by email

Daily job updates · Unsubscribe anytime