Jobs in United States

Software Engineer Ml Infrastructure Platform in United States

2,007 active opportunities · Updated October 2026

Explore current software engineer ml infrastructure platform jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

Hiring demand

51/100

steady · 543 related jobs

Hiring trend

-76.6%

Job postings compared with the previous 30 days

Remote options

15.1%

Share of matching jobs listed as remote

Typical salary

$177.2K – $177.2K/yr

Based on 29 salary observations

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -79.2%

$177.2K – $208.6K/yr · Jobiba est.

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the machine learning infrastructure that powers OpenAI’s monetization and ads systems. In this foundational role, you’ll design and develop the platform layer that enables teams to build, train, deploy, serve, monitor, and continuously improve machine learning models used across advertising and monetization products. You’ll work across the full ML lifecycle, from large-scale data pipelines and feature infrastructure to training systems, model serving, experimentation platforms, and monitoring frameworks. The systems you build will support high-throughput, low-latency advertising workloads while maintaining strict standards for reliability, privacy, security, and performance. This role sits at the intersection of machine learning systems, distributed infrastructure, and monetization, offering the opportunity to shape the core platforms that help translate model innovation into m

AWSRestMachine LearningAI
S
📍 Bellevue, Washington, United States· Full-time· Remote
✓ High-confidence listingCompany trend -91.7%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

KubernetesAIGoRust
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -91.7%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

KubernetesAIGoRust
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions. As a Senior Software Engineer in our Data Infra org, you will design, build, and scale the distributed data infrastructure platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role combines high ambiguity and ownership to push the boundaries of what our infrastructure can handle at massive scale, giving you the unique opportunity to steer the evolution of the data landscape. You Will: Own and Scale Core Platform Components: Take responsibility for the design, architecture, and implementation of 1–2 key data platform frameworks within our stack Collaborate and Align: Partner with infra, data science, and product engineering teams to ensure your target platform's capabilities are directly guided by platform governance and product requirements. Optimize Performance at Scale: Dive deep into engine internals, query planning, state management, memory optimization, serialization efficiency to maximize throughput and reliability under heavy load. Drive Infrast

JavaAWSGCPKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -79.2%

$177.2K – $208.6K/yr · Jobiba est.

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b

PythonAWSRestAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $278.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. ML Platform @ Roblox today supports hundreds of ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As an Infrastructure Engineer on the ML Platform team, you will design, scale, and maintain the foundational infrastructure powering our entire machine learning ecosystem. We are looking for accomplished engineers to spearhead the development of our next-generation ML tooling and platform capabilities. You will: Bootstrap and maintain Kubernetes and Cloud infrastructure for ML Platform components--Serving Layer, Metadata Store, Model Registry, and Pipeline Orchestrator. Set technical strategy and oversee development of high scale and reliable infrastructure systems. Propose and implement new platform tooling to improve time to production for MLEs and Data Scientists across the full ML lifecycle. Work on infrastructure projects such as GPU fleet management, hybrid-cloud orchestration, and writing custom Kubernetes controllers and resources. Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices. Partner across organizations to build tooling, interfaces, and visualizati

AWSGCPDockerKubernetes
R
📍 New York City, NY, United States· Full-time
✓ High-confidence listingBelow typical payDemand 51/100Company trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Data Platform builds infrastructure and tools that enable Ramp to realize business value from data. We partner closely with stakeholder teams to build this infrastructure and the applications on top of it. This role is particularly focused on building platforms that support the data science development lifecycle. You’ll partner with applied scientists, AI engineers, Risk engineers, and other ML developers on building infrastructure and tools that enable and accelerate the development of machine learning models. What You’ll Do Build and integrate the components of Ramp's Analytics Platform and Machine Learning Platform. Build tools that improve the agility and data experience of Ramp's Applied Scientists, AI Engineers, and Risk Engineers. Collaborate with stakeholder teams on building and productionizing machine learning applications. Build reliable, scalable, maintainable, and cost-efficient systems across the stack. What You Need Experience with workflow orchestrators like Airflow, Dagster, or Prefect. Experience building infrast

PythonSQLAWSAzure
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -73.6%

$177.2K – $208.6K/yr · Jobiba est.

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Infrastructure Software Engineer at Baseten, you'll build and maintain components of our ML inference platform that powers production AI applications. You'll contribute to the core infrastructure, enabling developers to deploy, scale, and monitor ML models with high performance. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Develop infrastructure components for our ML inference platform using Python and Go Implement and maintain Kubernetes deployments for model serving Contribute to our inference orchestration layer for model deployments Build and enhance monitoring systems for model performance metrics Implement efficient resource management solutions for ML workloads Support infrastructure automation to improve ML deployment workflows Work closely with team members to implement technical solutions Help balance performance optimization with system reliability Participate in technical discussions around infrastructure improvements Learn and apply infrastructure best practices REQUIREMENTS Bachelor's degree or higher in Computer Science or related field Proficient coding abilities in one or more popular programming or scripting languages; Go proficiency is a plus Working knowledge of Kubernetes and containeriza

PythonKubernetesRestMachine Learning
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -73.6%

$177.2K – $208.6K/yr · Jobiba est.

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues. Apply and scale optimization techniques across a wide range of ML models, particularly large language models. Collaborate with a diverse team to design and implement innovative solutions. Own projects from idea to production. REQUIREMENTS Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field. Experience with one

PythonDockerKubernetesRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -79.2%

$177.2K – $208.6K/yr · Jobiba est.

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

AWSRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -73.6%

$177.2K – $208.6K/yr · Jobiba est.

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg

PythonAWSGCPKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -79.2%

$177.2K – $208.6K/yr · Jobiba est.

This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea

AWSAzureKubernetesCI/CD
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -70%

$177.2K – $208.6K/yr · Jobiba est.

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Making data driven decisions is key to Plaid's culture. To support that, we need to scale our data systems while maintaining correct and complete data. We provide tooling and guidance to teams across engineering, product, and business and help them explore our data quickly and safely to get the data insights they need, which ultimately helps Plaid serve our customers more effectively. Engineers on Data Infrastructure are domain experts in Data Warehouse, Data Lakehouse, Spark, Workflow Orchestration, and Streaming technologies. We scale our existing data pipelines in a performant and cost efficient way while creating the necessary abstractions to make developing on top of this platform extremely simple for other engineers at Plaid. Responsibilities Contribute towards the long-term technical roadmap for data-driven and machine learning iteration at Plaid Leading key data infrastructure projects such as improving ML development golden paths, implementing offline streaming solutions for data freshness, building net new ETL pipeline infrastructure, and evolving data warehouse or data lakehouse capabilities. Working with stakeholders in other teams and functions to define technical roadmaps for key backe

PythonAWSMachine LearningAI
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedDemand 51/100Company trend -91.7%

$177.2K – $208.6K/yr · Jobiba est.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build the future of data. Join the Snowflake team. The Snowflake Machine Learning Platform team’s mission is to enable customers to bring their machine learning and deep learning workloads to Snowflake. Our customers want to build powerful models with the ever-increasing data in Snowflake but face several challenges including infrastructure optimizations, orchestration, performance, and security. The team aims to solve these challenges by building highly integrated platform solutions that are simple, secure, and enable end-to-end ML workflows. We are on an early journey to build the most scalable machine learning and data platform without sacrificing the benefits of a single platform and governance. We are looking for outstanding technical leaders who will join our ML Platform team to build the next-generation platform and play a pivotal role in this journey by understanding Snowflake’s core platform architecture and evolving it to enable state-of-the-art machine learning and LLM workloads. Join us to define strategies, set technical directions, design and execute, engage and deliver innovation, and unlock the power of AI for thousands of enterprise customers. This position is based in Menlo Park, CA, and Bellevue, WA. RESPONSIBILITIES : Help define and own the roadmap, wor

VueMachine LearningAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingTop 10% payDemand 51/100Company trend -100%

From $295.3K/yr

Quick readTop 10% pay versus similar roles

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions for a platform serving 200M+ daily active users. As a Principal Software Engineer in our Data Infra org, you will be the primary technical leader driving the strategic vision, long-term architecture, and massive scalability of our distributed data platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role operates under high ambiguity, demanding unparalleled ownership to redefine the limits of infrastructure handling exabyte-scale workloads, and providing a unique opportunity to lead the future evolution of our global data ecosystem. You Will: Define Multi-Year Technical Strategy: Own and drive the end-to-end architectural vision for Roblox's core data platforms spanning Kafka, Flink, Spark, Trino, Druid, Airflow, and Data Catalog systems. Turn multi-year company strategies into concrete, production-grade infrastructure blueprints. Lead Cross-Functional Alignment: Partner closely with executive leadership, platform governance, data science, and product e

JavaAWSGCPKubernetes

Related career options

Similar roles with stronger pay

Client Service Associate

Demand 46/100 · 8 jobs

$840K – $840K/yr

Salary →

$840K – $840K/yr

Salary →
Director of Product

Demand 43/100 · 6 jobs

$382.5K – $382.5K/yr

Salary →
Physical Design Engineer

Demand 43/100 · 8 jobs

$300K – $300K/yr

Salary →
Sr. Engineer

Demand 42/100 · 7 jobs

$300K – $300K/yr

Salary →
Senior Director

Demand 43/100 · 22 jobs

$278.9K – $278.9K/yr

Salary →
🔔

Get new software engineer ml infrastructure platform jobs in United States by email

Daily job updates · Unsubscribe anytime