About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
Jobiba hiring network
Senior Infrastructure Engineer Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Development Infrastructure team builds the tooling and systems our Asana engineers use every day to bring their ideas to production quickly and reliably. We build and operate the software that drives Asana’s roadmap. Each day, we combine industry best practices and innovation to support this product-focused company. We’re looking for an experienced Software Engineer with a passion for developer infrastructure. You will work with a world-class team of engineers on deploying and operating existing developer tooling, and building new tools to support our global, growing development team. You will have a unique opportunity to design and develop the systems and applications that drive the Asana development experience, lead complex technical projects, and work on cross-functional initiatives to help define the future of software engineering at Asana. This role is based in our Reykjavík office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Innovate the architecture of our development sandboxes to accelerate development iteration. Modernize our build system, tighten iteration loops, and polish frequent cycles in the developer experience. Lead complex developer infrastructure projects from technical design through implementation, rollout, and operation. Partner with engineering teams to identify opportunities to enable teams to develop faster at Asana and help the company achieve our goals faster. Analyze complex developer infrastructure systems to uncover issues, root causes, and areas for improvement. Keep Asana up to date on open-source trends and identify new opportunities for improved development infrastructure. Champion co
About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We're entering a world where AI agents don't just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents' ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova's throughput, correctness, and operational rigor grows dramatically. We're looking for a Senior Software Engineer who wants to go deep on the engine internals and the infrastructure underneath. You'll own significant components of a modern OLAP system — across query execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — and drive meaningful improvements to performance, cost-efficiency, and reliability. You'll grow your technical influence through the quality of your code, your design contributions, and your collaboration with other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of enterprise customers. What You'll Do Build and improve core query engine components Contribute across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Help ensure Nova's components support
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. The opportunity Global trade runs on data. Every shipment, every customs clearance, every delivery depends on the right information flowing through the right systems at the right time. At Flexport, data infrastructure is the foundation that enables billions of dollars in global commerce. We're looking for a Senior Software Engineer, Backend who will take end-to-end ownership of the data systems that power Flexport's global trade platform. In this role, you'll own critical data pipelines and architecture that process millions of supply chain events daily, ensuring that our customers have real-time visibility into their shipments across the world. As a Senior Software Engineer focused on Data Infrastructure and Backend development, you'll have deep technical ownership over the systems that ingest, process, store, and deliver data across Flexport. You'll architect solutions that handle the complexity of global trade, from tracking containers across oceans to coordinating multi-modal logistics, all while ensuring our data infrastructure scales reliably as we grow. You will Build strong relationships with producers and consumers of data across the organization. Partner
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will help develop and validate next-generation server platforms that power a reliable, high-performing, and cost-efficient infrastructure at scale. You will work across platform bring-up, firmware qualification, hardware validation, fleet integration, and performance optimization to support large-scale production deployments. You Will: Bring-up & Sustaining: Drive key aspects of the hardware development lifecycle, including feasibility studies, hardware bring-up, validation, deployment, and ongoing production support. Platform Optimization: Perform platform integration, performance characterization, and system-level debugging across compute infrastructure, focusing on hardware optimization, driver tuning, and thermal/power efficiency. Hardware Validation: Develop and execute rigorous evaluation and stress-testing strategies for server platforms to ensure reliability and performance under production-scale workloads. Firmware & Fleet Enablement: Support BIOS/BMC firmware qualification, hardware health monitoring, and automation tooling for firmware deployment and lifecycle management. Vendor & Cross-Functi
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We are a team of high-output generalists where ML and systems engineering converge to push autonomy performance forward. As a Senior Perception ML Data Infrastructure Engineer, you will own the critical bridge between our autonomous vehicle hardware, our human labeling operations, and our ML models. You will take ownership of our core native perception data platform. This is not a standard web development role; the stack is deeply adjacent to our core robotics infrastructure. You will be dealing with massive, dense 3D point clouds at a scale that pushes the boundaries of industry state-of-the-art, alongside complex sensor parsing, pipeline propagation, and rigid performance constraints. You will operate in highly ambiguous environments, inheriting complex systems, establishing strict API boundaries, and building the "good enough, fast enough" infrastructure that guarantees our ML models learn from high-quality data. About the wo
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs. The InfraSec team collaborates closely with other engineering teams to ensure that our infrastructure adheres to the highest security standards. They build essential security infrastructure and implement controls that reinforce the platform’s security posture. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions.This team is deeply involved in the technical aspects of security and the nuances of its actual implementation. This role can sit in our New York City, Austin, Seattle or San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. Responsibilities: Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security posture management (CSPM) Automation and Monitoring: Build automated solutions for real-time security monitoring, logging, and alerting in cloud environments. Leverage native cloud services and third-party tools for runtime security monitoring and anomaly detection Security Tooling: Evaluate, implement, and manage cloud-native security tools and platforms for endpoint security, identity management (IAM), and CSPM Qualifications: Experience: 6+ years of experience in SRE, infrastructure engineering or similar role, with a strong focus on security work, with ideally 2+ years in a senior or staff engineering role Security Mindset: A comprehensive understanding of all facets of cloud environment security, spanning from foundational OS networking laye
About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
Role: Senior Cloud Engineer Location: Hyderabad, India (Hybrid) Department: Product Development About GHX: GHX (Global Healthcare Exchange) is a leading healthcare technology company on a mission to simplify the business of healthcare and improve patient outcomes. Founded in 2000, GHX has built the GHX Global Network — the world’s largest cloud-based supply chain community connecting healthcare providers, suppliers, distributors, and partners to automate key processes, reduce costs, and increase operational efficiency. Its solutions span electronic trading, procurement automation, inventory and contract management, business intelligence, and data synchronization, helping healthcare organizations improve productivity and focus more on patient care. Over the years, GHX has enabled significant cost savings for the industry and continues to innovate with intelligent automation and AI-driven capabilities. Website: https://www.ghx.com/ LinkedIn: https://www.linkedin.com/company/ghx/ About Role: The Senior Cloud Engineer leads the design, implementation, and operations of the organization’s cloud infrastructure. This role is responsible for complex projects, high-level architectural planning, and making strategic technology decisions that support scalability, performance, security, and cost optimization. Acting as a technical leader, the Senior Cloud Engineer provides mentorship to junior and mid-level engineers while driving innovation, resiliency, and compliance in cloud environments. Key Responsibilities: Design and Implementation Lead the design and implementation of cloud-native architectures and hybrid cloud solutions. Oversee cloud engineering projects, ensuring solutions are secure, resilient, and cost-effective. Contribute to high-level architectural planning and strategic decision-making around technology adoption. Review and approve design proposals, Infrastructure as Code (IaC) tem
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and
We're looking for a Senior Software Engineer to integrate and improve our AI developer experience so that engineers at Asana can use AI to increase their velocity. As part of the AI Developer Productivity team, you'll build the next generation of AI-powered developer tools across editors, IDEs, CLIs, code review, and cloud and local coding agents.This role is based in our New York City office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Design and build AI-augmented workflows and systems that help coding agents understand the codebase, follow engineering practices, and enable Asana engineers to complete software development tasks faster and with more confidence. Design and scale autonomous cloud agents that take on complex, multi-step engineering tasks to reduce toil and enable engineers to focus on higher-leverage work. Build and refine IDE, editor, and CLI integrations that make AI-assisted development feel intuitive in engineers' daily work. Create reusable agent skills, tools, context, and integrations that teams across Asana can build on rather than reinvent. Improve AI-assisted code review workflows so engineers get faster, higher-quality feedback before and during review. Drive adoption of AI developer tools across engineering through usability improvements, measurement, documentation, and enablement. About you 5+ years of working in large codebases Experience building developer tools, internal platforms, infrastructure, IDE/editor integrations, CLIs, or workflow automation. Hands-on experience with AI-assisted development tools (such as Cursor, Claude Code, or Codex), or extending coding agents, MCP-based int
Asana’s rapid growth brings new challenges in keeping our systems fast, reliable, and resilient. As our product evolves, we’re making a major investment in reliability – and building a brand new SRE team in Warsaw is a key part of that strategy. This is your chance to help shape it from day one. This isn’t a traditional “ops” role – we’re looking for strong software engineers who are passionate about building reliable, distributed systems. You’ll work closely with a small SRE team in San Francisco, infrastructure engineers in Reykjavik, and an established infrastructure team in Warsaw. Warsaw will be a significant hub for our future infrastructure engineering and operations. As one of the first engineers here, you’ll have a real say in how we build reliable infrastructure, manage incidents, and support the rest of the company. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve Influence the future of Asana’s SRE practice, especially as we grow the Warsaw team. Lead reliability-focused projects across our stack – from infrastructure to tooling to incident response. Define and implement Asana’s incident management process – we’re investing here, and you’ll help shape how it works. Build internal platforms and frameworks that help other teams improve the reliability of their services. Be part of (and help shape) a sustainable on-call rotation – shared across teams in Warsaw, San Francisco, and Reykjavik. On average, we handle ~1 page per day, but it’s not constant, and we care about keeping things sane. Work with our stack: AWS, Kubernetes (EKS), Datadog, MySQL (RDS), ElasticSearch (OpenSearch), Redis
We're looking for a Senior Platform Reliability Engineer who brings strong software engineering skills and a deep understanding of system behavior under load and stress. This role is a good fit for someone who wants to own reliability as a first-class concern – building the foundational systems that protect Asana's platform, not just responding when things go wrong. You'll build core platform systems like load shedding, rate limiting, circuit breakers, and traffic controls that protect Asana under real-world load. This is deep, cross-cutting work that shapes stability and performance of our entire infrastructure – and you'll partner closely with other platform teams to make reliability something that's built in, not bolted on. Our tech stack includes: AWS, Kubernetes (EKS), CloudFront, Istio, Cilium, MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. (Yeah, we know this sounds like buzzword bingo – but we want this post to actually show up in your searches.) Why this role? Reliability as a first-class feature : You won't be patching things up after the fact. You'll build the systems that make Asana resilient by design. Foundational work : Load shedding, traffic management, ingress/egress – these are the building blocks that protect everything else. You'll own them. Strong collaboration, reasonable hours : You'll work closely with infrastructure teams in Warsaw and Reykjavik, making deep collaboration practical without constant timezone gymnastics. Room to grow : This is a new team, and you'll help shape what Platform Reliability Engineering looks like at Asana – whether that means leading projects, mentoring others, or defining our technical direction. In this role, success means shipping systems that other teams rely on by default – because they make the platform safer, not because they're mandatory. We're especially interested in people who think like backend engineers but obsess over failure modes, capacity plan
We’re looking for an experienced backend engineer with a passion for learning and working on systems. You will work with a world-class team of engineers on deploying and operating existing systems, and building new ones for challenges that are unique to our problem space. You will have a unique opportunity to design, develop, and operate services and frameworks that power Asana. This role is based in our Reykjavik office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Design, build, and iterate on frameworks that influence all aspects of the Asana product. Partner with other frameworks and product engineering teams to identify opportunities to enable teams to develop faster at Asana, and help the company achieve our goals faster. Analyze complex systems to uncover issues, root causes and areas of improvement Champion code quality and best practices, setting a standard through example and frameworks Develop new APIs that balance usability with the capability to support Asana’s feature-rich web application Write clean, beautiful code, striving to leave it in a better state than you found it Experience growth through opportunities to stretch and learn, enhancing your development Collaborate with a world-class engineering team to build new functionality through intuitive building blocks. About you 5+ years of BE experience working in large, well-maintained codebases writing and shipping production code Ability to learn quickly and transition seamlessly between different areas of a complex codebase Sound autonomous judgment when balancing moving quickly with producing quality, long-term maintainable code Focus on establishing and spread
Get new senior infrastructure engineer jobs by email
Daily job updates · Unsubscribe anytime