Jobiba hiring network

Software Engineer Infrastructure Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software engineer infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. In this role, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and deliver solutions that ad

typescriptjavareact
View job →
PE
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Are you passionate about engineering quality, performance, and increasing the impact of engineers around you? Software Engineers at Palantir build software at scale to transform how organizations around the world use data. As an engineer within Palantir’s Foundations organization, you’ll have the opportunity to grow more quickly than you ever imagined, as you build the shared infrastructure that underpins the Palantir Foundry, Palantir Gotham, and Palantir Apollo platforms, and drive investments to improve the velocity and quality of our engineering. Teams within Palantir’s Foundations organization are made up of a small number of engineers, each focused on one of four major categories of our infrastructure: • Backend Infrastructure: Maximizes the productivity of our backend developers and ensures Palantir’s platforms have performant and consistent RESTful services. Think: making the “easy way” the “right way” when developing backend services, including designing infrastructure to build hundreds of micro-service repos performantly, or to ensure we keep reliable audit logs of everything users do in our platforms. • Developer Infrastructure: Operates the systems and services that underpin all aspects of our developer ecosystem, including off-the-shelf tooling like GitHub and custom tooling for managing automated changes across hundreds of repositories. • Frontend Infrastructure: Maximizes frontend developer productivity across the entire frontend development stack, from the developer experience in the IDE to the final user experience in the browser. Think: the core infrastructure required to develop and serve our frontends (including features flags, int

typescriptjavareact
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. As an engineer within Palantir's Infrastructure teams, you'll have the opportunity to grow more quickly than you ever imagined as you contribute high-quality code directly to: The shared infrastructure underpinning Palantir Foundry, Palantir Gotham and Palantir Apollo — platforms deployed at the most important institutions across the public and private sectors Rubix and Mission Manager, our new internal-infrastructure business line, used by advanced civil and defence agencies worldwide to power their infrastructure in highly sensitive environments The substrate on which Palantir deploys Foundry and Gotham, powering workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters This means driving investments that improve the velocity and quality of our engineering. Infrastructure at Palantir spans our Foundations, Production Infrastructure and Foundry teams. Teams within Palantir's Foundations organisation are made up of a small number of engineers, each focused on one of four major categories of our infrastructure: Backend Infrastructure Developer Infrastructure Frontend Infrastructure Storage Infrastructure Production Infrastructure organisation, made up of small teams of engineers working on: Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters Apollo: secure, fleet-wide deployment and change-management for complex microservice suites Signals: our full suite of observability and alerting tools Foundry itself is also a developer

kubernetesaigo
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. In this role, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and deliver solutions that ad

typescriptjavareact
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. As a Software Engineer Intern, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors. • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters. You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer Intern at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and

typescriptjavareact
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Are you passionate about engineering quality, performance, and increasing the impact of engineers around you? Software Engineers at Palantir build software at scale to transform how organizations around the world use data. As an intern within Palantir’s Foundations organization, you’ll have the opportunity to grow more quickly than you ever imagined, as you build the shared infrastructure that underpins the Palantir Foundry, Palantir Gotham, and Palantir Apollo platforms, and drive investments to improve the velocity and quality of our engineering. Teams within Palantir’s Foundations organization are made up of a small number of engineers, each focused on one of four major categories of our infrastructure: • Backend Infrastructure: Maximizes the productivity of our backend developers and ensures Palantir’s platforms have performant and consistent RESTful services. Think: making the “easy way” the “right way” when developing backend services, including designing infrastructure to build hundreds of micro-service repos performantly, or to ensure we keep reliable audit logs of everything users do in our platforms. • Developer Infrastructure: Operates the systems and services that underpin all aspects of our developer ecosystem, including off-the-shelf tooling like GitHub and custom tooling for managing automated changes across hundreds of repositories. • Frontend Infrastructure: Maximizes frontend developer productivity across the entire frontend development stack, from the developer experience in the IDE to the final user experience in the browser. Think: the core infrastructure required to develop and serve our frontends (including features flags, inter

typescriptjavareact
View job →
N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The Autonomy ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving robots! About the Work Optimize Nuro’s autonomy stack with cutting-edge optimization techniques like quantization, low precision inference, and model pruning. Work with autonomy engineers to optimize, validate, and deploy large language models. Develop and maintain a world-class model compiler framework, FTL . Write robust, high-quality software to increase our confidence in our vehicl

pythonmachine learningai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team Frontier models are fungible. Any team can rent the same intelligence we can, and the model we build on today will be replaced within a month. What is not fungible is the infrastructure that decides whether an autonomous system's output can be trusted — evaluation, verification, and the discipline to gate on evidence instead of impressions. Nuro has spent a decade building exactly that discipline for a robot that drives on public roads, and this team turns it inward: we build the platform that lets AI agents operate autonomously inside Nuro's own engineering organization, under the same standard of proof we apply to the vehicle. Our mandate is to amplify the output of every engineer and researcher at Nuro by 100x. Not a better IDE, not a faster build — a change in what a single person can attempt. That number is a target, not a claim, and reaching it depends on one thing above all: autonomous work has to be trustworthy eno

pythonaic++
View job →

About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int

pythonawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

awsrestmachine learning
View job →
O
1mo ago

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the machine learning infrastructure that powers OpenAI’s monetization and ads systems. In this foundational role, you’ll design and develop the platform layer that enables teams to build, train, deploy, serve, monitor, and continuously improve machine learning models used across advertising and monetization products. You’ll work across the full ML lifecycle, from large-scale data pipelines and feature infrastructure to training systems, model serving, experimentation platforms, and monitoring frameworks. The systems you build will support high-throughput, low-latency advertising workloads while maintaining strict standards for reliability, privacy, security, and performance. This role sits at the intersection of machine learning systems, distributed infrastructure, and monetization, offering the opportunity to shape the core platforms that help translate model innovation into m

awsrestmachine learning
View job →
O
1mo ago

About the Team Our London-based team builds the backend systems that help ChatGPT scale reliably. We work on infrastructure close to the product, partnering with engineering teams to improve the performance, resilience, and operability of critical user-facing systems. Our work combines backend software engineering with distributed systems and production reliability. We build shared capabilities, improve high-traffic workflows, and make it easier to introduce new product functionality without compromising performance or availability. About the Role This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability. You’ll also own how those systems behave in production: how changes are rolled out, how issues are detected and diagnosed, and how recurring operational problems can be addressed through better software and system design. This is a strong fit for backend engineers who enjoy complex systems problems and want a direct connection between the infrastructure they build and the experience of ChatGPT users. In this role, you will: Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences. Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely. Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow. Build and improve systems for asynchronous processing and other large-scale backend workloads. Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on-call, incident response, and root-cause analysis, and turn operational lea

awsrestai
View job →
O
1mo ago

About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and

awskubernetesrest
View job →
O
1mo ago

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati

pythonsqlaws
View job →
O
1mo ago

About the Team At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products. Behind the scenes, making these systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for building a caching layer that powers many critical use cases at OpenAI. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis, Memcached), networking fundamentals, and Kubernetes-based service orchestration. In This Role, You Will: Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences. Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost. Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs. You Might Thrive In This Role If You: Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems. Have deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning. Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems. Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities. Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence. About OpenAI OpenAI is an AI research and deployment company d

redisawskubernetes
View job →
🔔

Get new software engineer infrastructure jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More software engineer infrastructure opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Software Engineer Infrastructure

Country links use the same curated canonical inventory as Jobiba sitemaps.