Jobiba hiring network

Staff Infrastructure Engineer Jobs

3,518 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current staff infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

A
Affirm
📍 Spain• Full-time• Remote• From €1.2M/yr
15 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Deal Reporting team is responsible for building our critical capital partner integrations, as well as services to automate funding processes safely and reliably. We partner with Product and Capitals Markets teams to understand and implement complex financial structures, integrations and reporting. We are looking for a highly motivated Staff Software Engineer to help empower Deal Reporting's first international team, and build integrations, services, and testing infrastructure to power the funding of every Affirm loan. What You'll Do: · You will be responsible for setting technical strategy for your team on a year-long time scale, and help your team tie it together with critical, business-impacting projects. · You will collaborate across teams in the product development lifecycle by collaborating with product management, design & analytics to ensure technical sustainability, risks and trade-offs are well understood and managed. · You will act as a force-multiplier for your team through your definition and advocacy of technical solutions and operational processes. · You take ownership of your team’s operations and availability by ensuring you have the right monitoring, triage rotations, playbooks, polcities, testing and alerting in place to support “keep the lights on” & on-call efforts. · You will foster a culture of quality and ownership on your team by setting code review and design standards for your team, and advocating for them beyond your team through your writing and tech talks. · You will help develop talent on your team by providing feedback and guidance, and leading by example. · “On-Call Rotation - There would be an on-call rotation for this role as a requirement”. What We Look For: · You have 7+ years of experience designing, developing and launching backend systems a

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Poland• Full-time• Remote• $576K – $864K/yr
15 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. The Deal Reporting team is responsible for building our critical capital partner integrations, as well as services to automate funding processes safely and reliably. We partner with Product and Capitals Markets teams to understand and implement complex financial structures, integrations and reporting. We are looking for a highly motivated Staff Software Engineer to help empower Deal Reporting's first international team, and build integrations, services, and testing infrastructure to power the funding of every Affirm loan. What You'll Do: · You will be responsible for setting technical strategy for your team on a year-long time scale, and help your team tie it together with critical, business-impacting projects. · You will collaborate across teams in the product development lifecycle by collaborating with product management, design & analytics to ensure technical sustainability, risks and trade-offs are well understood and managed. · You will act as a force-multiplier for your team through your definition and advocacy of technical solutions and operational processes. · You take ownership of your team’s operations and availability by ensuring you have the right monitoring, triage rotations, playbooks, polcities, testing and alerting in place to support “keep the lights on” & on-call efforts. · You will foster a culture of quality and ownership on your team by setting code review and design standards for your team, and advocating for them beyond your team through your writing and tech talks. · You will help develop talent on your team by providing feedback and guidance, and leading by example. What We Look For: · You have 7+ years of experience designing, developing and launching backend systems at scale using languages like Python or Kotlin. · You have an extensive track record of dev

REMOTEpythonsqlmysql
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $240K – $270K/yr
15 days ago

About the Role Sigma is transforming how businesses run by delivering a high performance platform on the modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing Solve challenging problems that arise in providing high performance interactive experience to enable analytics and workflows use cases on top of modern warehouses Build software using the latest developer tools and using programming languages like Rust, Go, GraphQL, Typescript Develop new algorithms and techniques for improving the performance and interactivity for enabling analytics and workflows for the world largest companies Triage product or system issues and debug/track/resolve by analyzing the sources of issues Design and implement new software features to support our fast growing user growth Collaborate with cross-functional groups - infrastructure, design, product, customer support, sales and marketing to build an innovative product capabilities Qualifications We Need 10+ years industry experience building and maintaining high-quality software Experience building and deploying robust and secure web applications in a continuous deployment environment Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong Computer Science fundamentals Qualifications We Want (also, skills you’ll learn!) Experience building software capabilities for analyzing large scale data web applications Data driven aptitude and its application to solve distributed system problems Data model design, and API development experience to enable customer facing cap

typescriptpythonsql
View job →
A
15 days ago

Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are currently seeking a Staff Software Engineer, Infrastructure to join the AI Platform team that powers seamless insights and interaction through natural language and data intelligence across our AI products. As a Staff Software Engineer, you’ll architect, build, and operate the backend and platform systems that power AI Platform. You’ll work across service design, distributed systems, cloud infrastructure, event-driven processing, observability, CI/CD, and production reliability, helping shape the technical direction of a platform that supports scalable, client-facing AI experiences. This role requires a strong software engineering foundation combined with deep infrastructure and systems thinking. We are looking for an engineer who can write high-quality production code, make sound architectural tradeoffs, and own platform capabilities end-to-end — not someone focused only on scripting, cloud configuration, or infrastructure tooling in isolation. You will collaborate closely with frontend, product, and AI/ML engineers to deliver reliable, secure, and scalable systems that align with Addepar’s standards of performance, resilience, and trust. Applicants must have legal authorization to work in the country where this role is based o

pythonjavasql
View job →
SA
15 days ago

About Scale Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust. About the Team Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities — and this role is not scoped to any single one of them. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like. About the Role As a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical needs — wherever the hardest ML problem in agentic AI happens to be that quarter. This could mean training and fine-tuning models, designing evaluation and observability systems, building improvement loops from production data, prototyping novel agent architectures, or designing internal systems and tooling that boost productivity across teams. You’re not tied to one team’s roadmap; you’re expected to move to where the technical leverage is highest, and t

awsrestmachine learning
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonaiexcel
View job →
G
Graphcore
📍 Austin• Full-time
15 days ago

About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced System Level Test Engineer to join our Product Test and Diagnosis Department (PTD). In this role, you will lead the development and deployment of System Level Test (SLT) solutions for next-generation AI processors. Working closely with cross-functional teams, you will contribute to the design and implementation of SLT hardware, software, automation, and characterization solutions that support silicon bring-up, yield learning, manufacturing readiness, and production deployment. The ideal candidate will possess strong technical depth in semiconductor test and validation, a passion for solving complex engineering challenges, and a strong focus on product quality and manufacturability. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Lead development and deployment of SLT hardware and software solutions supporting silicon bring-up, charac

pythonaisem
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track

pythonci/cdlinux
View job →
DU
15 days ago

About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and fi nancial reporting. Team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Sta ff Software Engineer,Data to be a technical lead and help architect and scale our data reliability, data infrastructure, automation and tools to meet growing business needs. You’re excited about this opportunity because you will... Own critical data systems that support multiple products/teams Develop, implement and enforce best practices for data infrastructure and automation Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Improve the reliability and scalability of our Ingestion, data processing, ETLs, Reporting tools and data ecosystem services Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We’re excited about you because... 8+ years of professional experience as a hands-on engineer and technical leader leading multiple projects 6+ years experience working in data platform and data engineering or a similar role You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Pro fi ciency in programming languages such as Python/Kotlin/Scala 4+ years of experience in ETL orchestration and work fl ow management tools like Air fl ow Expert in database fundamentals, SQL, data reliability practices and distributed computing 4+ years of experience with the Distributed data/similar ecosystem (Spark, Presto) and streaming technologies such as Kaa/Flink/Spark Streaming Excellent communication skills and experience working

pythonsqlaws
View job →
O
Okta
📍 Toronto• Full-time• From C$168K/yr
16 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About Okta Secures AI Team The Okta for AI Agents Team is building the future of digital identity management. The way companies are using agents and allowing them access is undergoing a fundamental shift, enabling teams to move fast while staying secure. Our North Star is clear: To get AI right, you have to get identity right. We are not just managing identities; we are building the industry's first Identity Security Fabric for the agentic era. As organizations rapidly deploy AI, these non-human entities are acting with access to critical tools, often bypassing traditional perimeters and creating a 'Shadow AI' crisis. Our mission is to move identity from a reactive gatekeeper to a unified control plane, transforming AI risk into business ROI. We are tackling this by implementing a comprehensive four-pillar maturity model: Discover, Onboard, Protect, and Govern. We are driving innovation by defining the standards for Agentic Identity—including pioneer work with securing AI Agents and support for protocols like MCP. You will be helping us build the infrastructure that allows enterprises to scale AI safely, ensuring every access decision—whether made by a human or an automated agent—is authenticated, authorized, and audited at scale. We are a diverse team of engineers, product managers, and designers who are bringing Okta’s expertise in identity to define the future of Agent Identity management, focusing on enablement while staying secure. About the role

typescriptpythonreact
View job →
O
Okta
📍 Bellevue• Full-time• From $194K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities—like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement b

pythonawsdocker
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What You’ll Do Work with various teams to design and implement scalable, and reliable network solutions Maintain a highly available cloud infrastructure edge for the Okta identity platform C

pythonawsdocker
View job →
O
Okta
📍 Australia; Sydney, Australia• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Identity Team Okta's internal Identity team secures the foundational trust layer for our modern enterprise. As organizations adopt phishing-resistant MFA, attackers have shifted toward session hijacking and cookie theft. Furthermore, with the rise of autonomous AI workloads, we are extending our platform to protect and govern a growing sprawl of non-human, agentic workflows. Our mission is to ensure every person and machine can safely connect to the right technology at the right time. About the Role As a Staff Engineer at Okta, you will be a technical authority within our internal IAM team, driving the architecture and execution of our identity fabric. You will champion our "Customer Zero" philosophy, partnering with product engineering to deploy Okta's newest features internally before they reach global markets. This role requires a technical leader who acts with ownership and urgency, turning action into traction and ruthlessly simplifying complex problems. You will serve as a mentor, influence architectural decisions, and help determine the future of the platform. Must be a U.S. citizen and will be working within the U.S. boundary to meet requirements of the FedRAMP compliance for Level 2. Here is the complete, consolidated job description tailored for the Staff Engineer level. What You’ll Be Doing (Responsibilities) Drive Customer Zero Execution: Act as a key technical bridge between internal stakeholders, the Okta on Okta team, and core Produ

awsazuregcp
View job →
O
Okta
📍 India• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Job Overview: We are looking for a highly experienced Staff Engineer to join the FGA DevEx team and lead the evolution of our end to end developer experience across both OSS and SaaS. This team owns the SDKs in Go, JavaScript, .NET, Python, Java and other languages, along with CLI workflows, IDE integrations, GitHub automation, developer documentation, and release strategy. All development is done in the open as open source, and we actively welcome and review community contributions. Our guiding principle is One developer experience, many deployment models. As a Staff Engineer, you will define technical direction, ensure cross language consistency, influence API design in partnership with FGA Core, and raise the quality bar across all developer facing tooling. Responsibilities: Define and drive the technical direction for FGA DevEx , CLI, IDE integrations, and developer automation across OSS and SaaS. Lead architectural decisions for multi language SDKs in Go, JavaScript, .NET, Python, and Java, leveraging and evolving the SDK generator that forms the core of all clients. Own and evolve the SDK generator framework, templates, and wrapper patterns to ensure cross language consistency, configurability, and long term maintainability. Establish standards across SDKs for authentication flows such as client credentials, error mapping and handling, retry logic with appropriate rate limiting strategies, and method level configuration overrides. Ensure advanced SDK

javascripttypescriptpython
View job →
A
Amplitude
📍 Remote• Full-time• $198K – $299K/yr
1mo ago

Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Staff Platform Engineer, you'll set technical direction for the platform across teams, lead our highest-complexity and highest-leverage initiatives, and shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll operate across team boundaries — partnering with product engineering, fellow Staff+ engineers, and engineering leadership to make Kubernetes and cloud infrastructure effortless across the entire engineering org. You'll build the self-service automation, shared standards, and scalable AWS and GCP infrastructure that let dozens of product teams ship faster, safer, and with less cognitive load — and you'll multiply the engineers around you while you do it. Key Responsibilities Set technical direction — shape platform and domain-level technical strategy that improves developer experience, reliability, security, and cost, and lead the high-complexity, cross-cutting initiatives that deliver it with measurable impact for the organization. Drive clarity through ambiguity. Take on the most loosely-defined problems, validate the critical assumptions early, and create alignment with stakeholders across teams so others can move quickly and confidently — driving cross-team decisions to a timely close and escalating when needed. Build the AI-augmented platform. Design org-wide tooling, guardrails, and policy-as-code that help every engineer get more out of AI-assisted development — infra primitives an LLM can safely reason about and PR against, automated review, and standards that hold as AI changes how code gets written. Own Infrastructure-as-Code standards for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — setti

pythonawsazure
View job →
🔔

Get new staff infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime