Jobiba hiring network

Staff Infrastructure Engineer Jobs

3,518 active opportunities · Updated for October 2026

Fresh results

13 shown

Explore current staff infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The ML Infrastructure team is responsible for building & improving the core infrastructure for autonomy teams at Nuro. In this role, you will work closely with teams across Nuro, to design, build and deploy core infrastructure components in machine learning model life cycle, to push the autonomous future forward. You will have an opportunity to work across the full stack of machine learning solutions - from designing robust and scalable model & data pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving robots! About the Work Design and develop ML workflow pipelines to train, optimize, validate, and deploy Nuro autonomy models. Develop and maintain continuous testing and monitoring systems for core ML infrastructure components. Develop observability to track ML model lifecycles from data generation to on-road validation. Maintain an in-house ML inference platform to serv

pythonmachine learningai
View job →
P
Pinterest
📍 United States• Full-time• Remote• From $208.6K/yr
1mo ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Intro : Pinterest’s Service Configuration & Coordination team builds the platforms, APIs, and tooling that make dynamic configuration and service coordination safe, scalable, and reliable for every service across regions and clouds, serving as the central hub for cross-functional alignment and system orchestration. As a Senior Staff Software Engineer (IC17) , you’ll own the end-to-end technical strategy and long-term vision for our configuration and coordination infrastructure. You’ll define the future of these mission-critical systems, partnering closely with Traffic, Compute, Observability, and Cloud Architecture teams to deliver robust, high-performance solutions at scale. What you’ll do: Orchestrate the long-term technical vision for configuration ecosystems to ensure feature flags, ML settings, and experiments utilize unified pave

REMOTEawsci/cdrest
View job →
A
Airbnb
📍 - USA• Full-time• Remote• From $248K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Machine Learning and Artificial Intelligence are at the heart of the Airbnb product. From Trust to Payments, and from Customer Service to Marketing we rely on ML to ensure that guests and hosts have the best possible experience with Airbnb. The CS AI product team is responsible for driving CSxAI (Customer Support x Artificial Intelligence) initiatives by adopting the Generative AI technologies to enable an intelligent, scalable and exceptional service experience. The team develops and enhances various AI models, ML services and tools including LLM fine-tuning, alignment and optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning and guardrail for a wide range of applications in Airbnb. The Difference You Will Make: As a senior staff machine learning engineer, you will be responsible for fine-tuning state-of-the-art LLMs for diverse use cases while optimizing models for high-performance deployment on Airbnb’s ML Infrastructure. You will partner with product managers, software engineers, data scientists and operation teams to brainstorm, design and develop AI products such as AI Assistant, Autonomous agent, recommendation, travel planning, and many more products that make meaningful impacts in the world of travel. A Typical Day: Work with large scale structured and unstructured data; explore, experiment, build and continuously improve foundation models for Airbnb product, business and operational use cases. Create a multi-year tech roadmap that enables our team to stay on the leading edge of the rapidly evolving AI landscape and

REMOTEpythonmachine learningai
View job →
A
Airbnb
📍 United States• Full-time• Remote• From $248K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Our mission at Marketing Technology is to provide a state-of-the-art platform and measurement capabilities to enable marketing and product teams to engage with our customers effectively. Every day, millions of Airbnb customers are reached via our streamlined platform over comprehensive marketing channels such as Email, Push, SMS, Google Ads, Facebook Ads etc. We deliver substantial business impact by enabling numerous stakeholders across the company, including Marketing (Brand and Performance), Guest, Host, Airbnb Org, Policy, and more. The Difference You Will Make: As a Senior Staff Software Engineer on the Content Platform team, you'll set the technical direction for the org's transformation into an AI-driven system that turns a marketing brief into a fully realized campaign in minutes, not months. You'll own the technical strategy across a complex, multi-team surface area spanning a CMS platform, no-code authoring tools, rendering infrastructure, and an emerging agentic layer, all operating at a scale where the content produced is consumed by millions of Airbnb users daily. You will partner with engineering leadership, product, design, and marketing stakeholders across the company to define multi-year technical strategy, resolve the org's hardest cross-team technical challenges, and raise the bar for engineering excellence across Marketing Technology. A Typical Day: As a Senior Staff Software Engineer on Content Platform, you will: Define and drive technical vision and strategy for the Content Platform org, ensuring architectural decisions scale across teams and ali

REMOTEjavasqlredis
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Enterprise Identity team builds the identity foundation that enables organizations to adopt and use OpenAI products securely and reliably. The team owns the enterprise identity stack, including SSO, SCIM, tenant architecture, and identity capabilities across the enterprise admin experience and OpenAI's growing multi-product portfolio. About the Role We are looking for a hands-on senior technical leader to own the architecture and evolution of OpenAI's Enterprise Identity systems. You will set the long-term technical vision for the entire stack, establish shared identity primitives across products, and be accountable for systems that are foundational to our enterprise business. This role requires operating well beyond a single service or feature area. You will identify the most consequential architectural investments, align teams around durable solutions, and ensure our identity platform meets an exceptionally high bar for scale, availability, latency, and security. This role will be based in our San Francisco or Mountain View office. In this role, you will: Own the technical vision and architecture for the Enterprise Identity stack, including SSO, SCIM, tenant architecture, groups, permissions, and identity capabilities in enterprise administration surfaces. Lead the design and evolution of highly available, latency-sensitive identity systems serving a large and diverse global enterprise customer base. Establish common identity models and primitives that work consistently across OpenAI's products and enable the organization to scale. Set a high security bar by anticipating abuse cases, failure modes, and the long-term implications of new capabilities. Drive alignment across enterprise product, infrastructure, and security partners, resolving ambiguity and influencing roadmaps beyond the immediate team. Provide technical leadership to senior engineers and raise the quality of architecture and execution across the broader organization. You might thr

awsrestai
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! The integration team is responsible for developing and scaling machine learning algorithms and infrastructure for LLM post-training, with a focus on large-scale, distributed RL methods. We strive for excellence in both engineering and science by meticulously designing experiments and design docs. While tasks are assigned according to everyone’s expertise, there is a global team effort to write production code and support the team research efforts, depending on individual interests and organizational needs. In particular, this role aims to enhance the global quality of the post-training codebase by implementing new tools to ease and support research, optimizing post-training algorithms, and scaling distributed RL to unprecedented levels. Please Note: We have offices in London, Paris, Toronto, San Francisco, New York but we are also remote-friendly! Applicants for this role may work anywhere between UTC−06:00 and UTC+01:00. As a Member of Technical Staff, you will: Design and write high-performing and scalable software for training models. Develop new tools to support and accelerate research and LLM training. Coordinate with other

pythonkubernetesgit
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role Our robotics team is growing and we are looking for a Software Engineer to join our Sensor Data and Calibration team. We are searching for an engineer with robotics and machine learning expertise to develop synthetic sensor simulation models and algorithms. The ideal candidate has hands-on experience in the research, development, and implementation of machine learning methods (e.g., NeRF or Gaussian splatting) for generating synthetic sensor data (photorealistic images, realistic lidar and/or radar, etc.). About the Work Research, develop, and implement state-of-the-art synthetic sensor simulation methods Analyze and characterize the realism and utility of synthetic sensor data Answer critical questions about sensor data and autonomy performance Collaborate with stakeholders across autonomy, infrastructure, and systems teams on map needs and requirements Role is scoped as a Senior/Staff IC with the flexibility to grow into

pythonmachine learningai
View job →
M
Mongodb
📍 Boston; Miami; New York City; Pittsburgh; Raleigh; United States• Full-time• From $126K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Boston, New York City, Raleigh, Miami, Pittsburgh or remotely in the United States while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure g

pythonmongodbaws
View job →

The Team This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design & build complex systems, operate with autonomy and act as owner for everything you do. The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers. The ideal candidate should Have 5+ years of experience running critical systems at scale Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”) Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment A strong understanding of how to run a large scale Linux environment, including low level fundamentals Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python) Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc) Special Requirements: Be a US Citizen Expectations Participate in the development of a reliable and resilient multi-cloud platform that hosts business critic

pythonmongodbaws
View job →

The Team MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of either our Dublin or Cork office or remotely in Ireland. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Build for reliability, making services and infrastructure avail

pythonmongodbaws
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Toronto or Montreal office or remotely in the Canada while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineerin

pythonmongodbaws
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold builders and sharp problem-solvers who are wired to deliver great outcomes. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. The DevX team’s mission is to build and operate the core developer infrastructure at Robinhood. Our team owns and scales the systems that thousands of engineers rely on daily, partnering with software developers across the company to make development fast, reliable, and cost-efficient! As a Staff Software Developer, you will act as a technical leader for our build and developer infrastructure, driving the strategy and execution of the systems thousands engineers depend on every day. Your work will span our build systems, CI pipelines, and remote development environments, ensuring engineers can code, test, and build with speed, safety, and reliability at scale. In this role, you will collaborate with teams across Robinhood to eliminate developer friction and raise the bar for engineering productivity. This is a high-visibility leadership opportunity to shape our developer ecosystem and set new standards of engineering efficiency! This role is based in our Toronto, ON office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. What you’ll do Architect the long-te

pythonawsci/cd
View job →
E
15 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pipeline Engineering team at Everpure™ as a Software Engineer to build, own, and operationally scale the microservices and Temporal workflows driving continuous integration for FlashArray, FlashBlade, and Hyperscale. In this developer-first role based in Bangalore, you will design durable production services and agentic AI tools that directly reduce time-to-signal, optimize compute utilization, and absorb operational toil across our global engineering workforce. WHAT YOU'LL DO Architect Deterministic CI Workflows: Design and deploy production microservices and durable Temporal workflows that make pipeline execution resilient, resumable, and fully debuggable for enterprise storage platforms. Optimize Compute & Testbed Utilization: Extend in-house scheduling logic and placement algorithms across bare-metal hardware and VM fleets to minimize queue times and maximize infrastructure efficiency. Build Operational AI Agents: Engineer RAG pipelines and agentic workflows over failure logs, vector stores, and test metadata to automate root-cause analysis, flake classification, and developer triage. Drive Platform Observability & Ownership: Establish key platform metrics—including time-to-signal and pass/flake rates—while operating your services end-to-end to ensure long-term stability and eliminate recurring failure modes. Collaborate Across Product Teams: Partner with platform and product engineering group

pythonsqlaws
View job →
🔔

Get new staff infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime