Jobiba hiring network

Senior Infrastructure Software Engineer Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior infrastructure software engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

D
Dropbox
📍 Poland• Full-time• Remote
1mo ago

Role Description As an Infrastructure Engineer, your role will be crucial in shaping and constructing the robust systems that not only support our current flagship products but also lay the groundwork for the next wave of engineering innovations. From optimizing user experiences across various projects to ensuring seamless scalability and data integrity, you'll be at the forefront of shaping the technological backbone of our platform. Collaborating closely with cross-functional teams, you'll leverage your expertise to tackle audacious challenges and push the boundaries of what's possible. Your contributions will directly impact millions of users, as every line of code you write furthers our mission to revolutionize the way people work and collaborate. Join us in redefining the future, where your passion for building scalable, reliable systems will drive meaningful change on a global scale. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Build infrastructure capable of managing metadata for hundreds of billions of files, handling hundreds of petabytes of user data, and facilitating millions of concurrent connections. Lead the expansion of Dropbox's function as the data-fabric, connecting hundreds of millions of applications, devices, and services globally, while also driving initiatives to enhance interoperability and adaptability across diverse ecosystems. M easur e and optimiz e Dropbox's analytics platform to maintain its status as one of the most advanced in the industry for extracting meaningful insights from vast data volumes. Collaborat e with cross-functional teams to innovate and implement solutions that enhance the performance, reliability, and security of Dropbox's infrastructure, ensuring a seamless experience for users worldwide. Proactively identify new opportunities and drive imp

REMOTEpythonjavagit
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

We're looking for a Senior Infrastructure Engineer who brings strong software engineering skills and a deep understanding of production systems. This role is a good fit for someone who enjoys building systems that make infrastructure more scalable, reliable, and easy to operate – using code, not runbooks. You'll work with a highly collaborative team to design and build the internal platforms that power all of Asana, from product features to AI systems to offline analytics. Our tech stack includes: AWS, Kubernetes (EKS), MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. We’re especially interested in people who think like backend engineers but care deeply about systems – things like failure modes, operational cost, debuggability, and performance. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve: Design and build frameworks, tools, and services that improve the reliability, observability, and scalability of Asana’s infrastructure. Lead end-to-end projects, from scoping and design through to rollout, across multiple systems and teams. Improve the operability of stateful infrastructure like MySQL, OpenSearch, and DynamoDB – and help drive Asana’s long-term vision for storage reliability. Debug production issues across the stack. Yes, there’s an on-call rotation – but this isn’t a pager monkey role. You’re here to fix things properly and make sure they don’t break again. Partner with product teams to shape a service-oriented architecture that enables fast, reliable development. Share knowledge through code reviews, design discussions, and mentorship. Abou

typescriptpythonsql
View job →
G
22 days ago

About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support the software build and release process Deploy and maintain services with Kub

pythonawsdocker
View job →
HI
HP IQ
📍 San Francisco• Full-time• $162K – $288K/yr
22 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ’s System Software team enables on-device experiences to take full advantage of our hardware capabilities. We collaborate with internal and external partners in a high-leverage environment that enables us to spend a majority of our time developing solutions that are unique to our hardware, sensors, algorithms, and interaction models. If you enjoy solving complex, interdisciplinary problems with a world-class team, we'd love to hear from you! What You Might Do Learn what it's like to be a part of a world-class embedded software team building a first-of-its-kind product in a startup environment Responsible for system design and architecture Develop low-level driver and framework software in C and C++ Develop device-focused infrastructure software in Python Debug issues at the interface between hardware and software Optimize software for better performance and lower power consumption Collaborate in the software engineering process with documentation, testing, and code review Essential Qualifications 8+ years of experience in system software engineering and embedded platf

pythonredislinux
View job →

About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We're entering a world where AI agents don't just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents' ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova's throughput, correctness, and operational rigor grows dramatically. We're looking for a Senior Software Engineer who wants to go deep on the engine internals and the infrastructure underneath. You'll own significant components of a modern OLAP system — across query execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — and drive meaningful improvements to performance, cost-efficiency, and reliability. You'll grow your technical influence through the quality of your code, your design contributions, and your collaboration with other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of enterprise customers. What You'll Do Build and improve core query engine components Contribute across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Help ensure Nova's components support

pythonjavaredis
View job →

About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine

pythonkubernetesci/cd
View job →
R
Roblox
📍 San Mateo• Full-time• From $278.5K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. ML Platform @ Roblox today supports hundreds of ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As an Infrastructure Engineer on the ML Platform team, you will design, scale, and maintain the foundational infrastructure powering our entire machine learning ecosystem. We are looking for accomplished engineers to spearhead the development of our next-generation ML tooling and platform capabilities. You will: Bootstrap and maintain Kubernetes and Cloud infrastructure for ML Platform components--Serving Layer, Metadata Store, Model Registry, and Pipeline Orchestrator. Set technical strategy and oversee development of high scale and reliable infrastructure systems. Propose and implement new platform tooling to improve time to production for MLEs and Data Scientists across the full ML lifecycle. Work on infrastructure projects such as GPU fleet management, hybrid-cloud orchestration, and writing custom Kubernetes controllers and resources. Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices. Partner across organizations to build tooling, interfaces, and visualizati

awsgcpdocker
View job →
G
22 days ago

About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team An exciting opportunity to join a new team within the Software Operations group. The Build Engineering team is a new function within Software Infrastructure, which focuses on the overall process of building and integration of the Machine Learn ing S oftware S tack. You will work closely with the QA and development teams to get an understanding of how our ML SW stack is built, helping to ensure good build practices, and proving that the stack works together and is reproducible in secure, sandboxed environments. Responsibilities and Duties Developing our internal t

pythondockerlinux
View job →

About BlockTech BlockTech is a fast-paced algorithmic trading firm facilitating global cryptocurrency derivatives and spot trading while expanding into new markets. As we continue to grow rapidly, we are looking for a Software Engineer to join our Foundation team amid our exciting scale-up phase! You will Build & optimize: Design, develop, and maintain high-reliability, low-latency, and high-throughput foundational systems that enable our trading and technology teams to scale efficiently. Ingest & aggregate: Collect trading business data with minimal latency impact and ingest both public and private exchange information into our trading system. Store & stream: Develop and maintain infrastructure for real-time data aggregation and long-term storage, as well as our Kafka-based messaging systems. Collaborate & support: Work closely with multiple teams, assisting them in integrating with and making the most of our foundational systems. Innovate: Drive projects from concept to deployment with full ownership, and explore new tools, frameworks, and approaches to keep our infrastructure best-in-class. The Foundation team develops core software infrastructure (libraries, frameworks, and systems) for BlockTech, solving common problems and lending its expertise to enable other teams to stay focused on their respective domains. They own, develop, and configure a wide variety of critical, high-reliability software, ranging from low-level ultra-low-latency shared memory IPC libraries to high-throughput data buses and data aggregation systems including Kafka, NATS, PostgreSQL, and Iceberg. They work primarily in Rust, but also use Python and SQL. If you thrive on low-level problem-solving, building robust frameworks from scratch, and enabling others to move faster, this role is for you. What We're Looking For Essential: 5+ years of experience as a Software Engineer, with a strong focus on systems-level optimisation and awareness of hardware constraints Proficiency

pythonsqlpostgresql
View job →
R
Roblox
📍 San Mateo• Full-time• From $196.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox’s daily active users growing at a record pace, we are seeking a senior data infrastructure engineer to join our new Data Insights team. Our team owns the data tooling that empowers Roblox builders to independently make informed and timely data-driven decisions. As an engineer on the team, you’ll work on the platforms behind tools like Superset, Hex, and Python notebooks, which provide critical insights into the health of our business to users at every level of the company. We tackle diverse challenges in data engineering, infrastructure, and analytics, to deliver the insights our customers need. You will collaborate closely with engineers across our data ecosystem to shape the future of product analytics at Roblox. This role offers the chance to be a founding team member and help define both the technical direction and the long-term shape of the product area from the ground up. This role is a great fit for you if you are proficient in designing and scale robust data infrastructure and applications and have a zeal for developing inspiring, easily maintainable, and reusable code. Join our team and make a significant impact at Roblox. You Will: Architect and deliver a high-pe

typescriptpythonreact
View job →
G
Gitlab
📍 United Kingdom; Remote, United States• Full-time• Remote• From $126.4K/yr
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure. This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience

REMOTEawsgcpkubernetes
View job →

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu

pythonkubernetesci/cd
View job →
F
Fin
📍 Dublin• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? We’re looking for Senior+ AI Infrastructure Engineers to build the systems that train and serve Fin's next generation of AI products. Fin is an AI company that builds from the GPU all the way up to a user agent that resolves millions of customer service queries a month. You’ll join a small, highly technical team working at the cutting edge of modern AI infrastructure. The AI Infra team built the training pipelines and runs the inference for custom models like Fin Apex, which outperforms frontier models in customer service tasks, and is the foundation of the AI Group's full stack approach to AI. We’re particularly interested in engineers who have: A track record of working on model training or model inference at scale , or on low‑level GPU coding (e.g. CUDA, Triton). Experience with one is great, multiple is even better. What will I be doing? As a Senior AI Infrastructure Engineer focused on model training and inference, you will: Implement and scale training pipeli

pythonjavaaws
View job →
F
Fin
📍 England• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? We’re looking for Senior+ AI Infrastructure Engineers to build the systems that train and serve Fin's next generation of AI products. Fin is an AI company that builds from the GPU all the way up to a user agent that resolves millions of customer service queries a month. You’ll join a small, highly technical team working at the cutting edge of modern AI infrastructure. The AI Infra team built the training pipelines and runs the inference for custom models like Fin Apex, which outperforms frontier models in customer service tasks, and is the foundation of the AI Group's full stack approach to AI. We’re particularly interested in engineers who have: A track record of working on model training or model inference at scale , or on low‑level GPU coding (e.g. CUDA, Triton). Experience with one is great, multiple is even better. What will I be doing? As a Senior AI Infrastructure Engineer focused on model training and inference, you will: Implement and scale training pipeli

pythonjavaaws
View job →
F
Fin
📍 Berlin• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? We’re looking for Senior+ AI Infrastructure Engineers to build the systems that train and serve Fin's next generation of AI products. Fin is an AI company that builds from the GPU all the way up to a user agent that resolves millions of customer service queries a month. You’ll join a small, highly technical team working at the cutting edge of modern AI infrastructure. The AI Infra team built the training pipelines and runs the inference for custom models like Fin Apex, which outperforms frontier models in customer service tasks, and is the foundation of the AI Group's full stack approach to AI. We’re particularly interested in engineers who have: A track record of working on model training or model inference at scale , or on low‑level GPU coding (e.g. CUDA, Triton). Experience with one is great, multiple is even better. What will I be doing? As a Senior AI Infrastructure Engineer focused on model training and inference, you will: Implement and scale training pipeli

pythonjavaaws
View job →
🔔

Get new senior infrastructure software engineer jobs by email

Daily job updates · Unsubscribe anytime