ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
Jobiba hiring network
Lead Infrastructure Software Engineer Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead infrastructure software engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens
Join the Atlas Search team to design and develop the next generation of Semantic and Vector Search infrastructure. Atlas Search is a growing cloud service that allows users to execute complex search and vector search queries using the MongoDB Query Language. Our users can focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building a cloud-based distributed system responsible for the core components of search including data ingestion, performance, query language, query execution, for both relevance-based search and vector search. Our product is being adopted quickly and there are many interesting projects. This is a technical role where you will be responsible for the infrastructure and features enabling our at-scale cloud service powering vector and semantic search. We are looking to speak to candidates who are based in the San Francisco Bay Area for our hybrid working model. What You’ll Do Lead complex projects across the MongoDB ecosystem, for instance, development of a new Search deployment framework within the MongoDB managed cloud Set project level strategy, architect features, and lead projects to successful execution Identify, design, and implement features enhancing our reliability, performance, security and efficiency Perform code reviews with peers and make recommendations on how to improve our software development processes Influence and grow team members through active mentoring and leading by example What We Look For 5+ years experience in data management/search systems, ideally with a strong distributed systems and infrastructure background Experienced in the development and maintenance of concurrent, stateful services Eager to shape the technological direction of a complex system and have the ability to lead initiatives through collaboration with others Experienced in writing features, debugging and optimizing multithreaded applications written in Java Familiarity with LLM
Scale AI is seeking a highly skilled and motivated Software Engineer, Frontier AI Infrastructure to join our dynamic Public Sector Engineering team. As a part of this team, you will own the model inference layer - enabling state of the art models, debugging the latest AI tools, managing networking, debugging latency, and tracking pricing/usage metrics for AI models. You will lead technical discussions on the frontlines with cloud vendors and customers to deliver on critical contracts and to debug platform issues. You will also work upstream with Product to understand features before they break, moving us from "infra-only debugging" to proactive integration testing. You will: Design and implement secure scalable backend systems for Public Sector customers, leveraging Scale's modern and cloud-native AI infrastructure. Own services or systems and define their long-term health goals, while also improving the health of surrounding components Re-architect the stack to run in compliant or restrictive environments. This requires designing swappable components (auth, storage, logging) to meet government/security mandates without breaking the product. You will work with Product to build integration tests that catch issues early, shifting the focus from "infra-only debugging" to preventing failures upstream. Participate actively in customer engagements, working closely with stakeholders to understand requirements and deliver innovative solutions. Contribute to the platform roadmap and product strategy for Scale AI's Public Sector business, playing a key role in shaping the future direction of our offerings. Must have: At least an active secret clearance and the ability & willingness to up level to TS/SCI with CI Poly. This is a requirement and candidates will not be considered who do not hold at least a secret clearance Ideally you'd have: Full Stack Development: Proficiency in both front-end and back-end development, including experience with modern web develo
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: This is a Principal Product Engineering role focused on Money Infrastructure at Replit. You’ll work on the financial backbone that powers how Replit earns money, how builders earn money, and how Agents transact. This role sits at the intersection of engineering, product, and the business. The systems you build directly impact revenue, trust, and some of the most critical user journeys on the platform. Getting them right enables growth, experimentation, and global scale. Getting them wrong creates broken payments, confusing pricing, and lost trust. We’re looking for engineers who can design and scale reliable financial systems while translating complex monetary logic into intuitive, user-friendly experiences for both Replit customers and builders on the platform. We love folks who have a passion for monetizing innovation and being a part of the greater pricing story. You will: Lead the design, architecture, and implementation of Replit’s core money infrastructure, spanning pricing, billing, payments, and monetization. Own and scale the global order-to-cash foundation supporting credit-based subscriptions, usage-based billing, marketplaces, in-app payments, and commerce for Agents. Enable rapid pricing and packaging experimentation across the company by building flexible abstractions and APIs for new SKUs, plans, and monetization models. Build high-converting, localized payment experiences across geographies — thinking globally while enabling users to pay locally. Power builder monetization by creating payment rails for apps, Agents, subscriptions, and new monetization primitives. Partner closely on data specifications with finance, accounting, and data teams to produce accurate, auditable, and reliable f
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions for a platform serving 200M+ daily active users. As a Principal Software Engineer in our Data Infra org, you will be the primary technical leader driving the strategic vision, long-term architecture, and massive scalability of our distributed data platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role operates under high ambiguity, demanding unparalleled ownership to redefine the limits of infrastructure handling exabyte-scale workloads, and providing a unique opportunity to lead the future evolution of our global data ecosystem. You Will: Define Multi-Year Technical Strategy: Own and drive the end-to-end architectural vision for Roblox's core data platforms spanning Kafka, Flink, Spark, Trino, Druid, Airflow, and Data Catalog systems. Turn multi-year company strategies into concrete, production-grade infrastructure blueprints. Lead Cross-Functional Alignment: Partner closely with executive leadership, platform governance, data science, and product e
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our New York City, Austin, Seattle or San Francisco offices, or work fully remotely on standard East Coast business hours. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our Dublin office, or work fully remotely in Ireland. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations and processes A deep understanding of Linux and networking concepts,
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Service Infrastructure team is responsible for empowering teams to build highly reliable and performant services, backing our payment systems, fraud detection, and a multitude of other products. As an infrastructure team, you will build and expand the service frameworks to make complex systems easy to use, resilient, and scalable. We build powerful interfaces for engineers that depend on these systems—deployment, load balancers, web framework, databases, Kafka, and Kubernetes—while keeping them highly available and performant. We're looking for engineering leaders who drive the technical vision of Stripe's service infrastructure platform for thousands of engineers to use and build on top of, and thrive in a highly autonomous environment with many moving pieces. What you’ll do You will join as a Technical Lead for one of the most impactful teams at Stripe. You will lead a team of engineers, collaborate with infrastructure and product engineering orgs, and advance service-oriented architecture (SOA) adoption at Stripe. By collaborating with the team's technical leaders, you will ensure the software your team builds meets the needs of Stripe and its customers. We are a highly effective team that consistently delivers high-impact results while genuinely caring for one another. We expect you to bring your curiosity and critical thinking skills to this challenging domain. We are looking for individuals with a strong background in designing and
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust
About the Team Our London-based team builds the backend systems that help ChatGPT scale reliably. We work on infrastructure close to the product, partnering with engineering teams to improve the performance, resilience, and operability of critical user-facing systems. Our work combines backend software engineering with distributed systems and production reliability. We build shared capabilities, improve high-traffic workflows, and make it easier to introduce new product functionality without compromising performance or availability. About the Role This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability. You’ll also own how those systems behave in production: how changes are rolled out, how issues are detected and diagnosed, and how recurring operational problems can be addressed through better software and system design. This is a strong fit for backend engineers who enjoy complex systems problems and want a direct connection between the infrastructure they build and the experience of ChatGPT users. In this role, you will: Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences. Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely. Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow. Build and improve systems for asynchronous processing and other large-scale backend workloads. Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on-call, incident response, and root-cause analysis, and turn operational lea
About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion’s Data Foundations team builds and operates the batch and streaming infrastructure behind our product features, analytics, search, and AI experiences. We’re looking for a hands-on technical leader to shape the next generation of this platform as Notion serves larger customers, expands globally, and supports more data-intensive products. You’ll identify the highest-leverage problems, set direction, build and develop a high-performing team, and lead multi-quarter initiatives across our data lake, streaming, distributed-compute, governance, and reliability systems. You’ll stay close to critical technical decisions while creating clear ownership, growing engineers and technical leaders, and helping the team execute as one—partnering closely with Data Engineering, Data Product, Search, AI, Infrastructure, and Security. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Data Platform organization builds and operates the systems that power how data is stored, moved, and consumed across Robinhood. This organization spans three core pillars: Storage (Postgres, DynamoDB, and caching systems), Streaming (real-time event infrastructure), and Data Lake (ingestion and compute systems built on Delta Lake). Together, these platforms support transactional workloads, real-time data processing, and large-scale analytics that are critical to Robinhood’s products and operations. The team owns the full lifecycle of data—from low-latency order path systems to near real-time and batch analytics—serving millions of users and internal teams across the company. As a Senior Staff Software Engineer , you will serve as the technical lead across the Data Platform organization, shaping architecture and guiding execution across multiple teams. You’ll work on complex distributed systems challenges such as database sharding and proxy architectures, real-time streaming and CDC systems, and large-scale data ingestion and compute platforms. You’ll define and drive key technical bets, partner with engineering leaders to align platform capabilities with business needs,
Get new lead infrastructure software engineer jobs by email
Daily job updates · Unsubscribe anytime