Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking a highly technical Lead Release Engineer with a strong engineering foundation to lead the end-to-end release lifecycle of our storage products. You will serve as the bridge between Development, QA, and Product Management — ensuring that complex storage stacks are delivered with high quality and predictable cadences. Unlike traditional project-based release management, this role demands deep hands-on expertise across CI/CD orchestration, codeline management, system-level triaging, fleet operations, and the engineering rigor required for data-critical products. You will own the health of our release pipelines, lead triage war rooms, drive automation initiatives, and participate in on-call rotations to keep CI and test-orchestration infrastructure running reliably. You will also build developer-facing tooling, manage HW test fleet operations, and maintain high-quality integration workflows across our code lines. WHAT YOU'LL DO Release Orchestration: Own the end-to-end release process for storage software and firmware, from development to GA (General Availability). CI/CD Leadership: Design and build optimized pipelines and tools to scale code management and merge operations. Work closely with systems such as Jenkins, test frameworks, Premerge, Orchestrator, and related developer productivity tooling to keep the codeline healthy and actionable. Technical Triaging: Act as the primary technical poin
Jobiba hiring network
Staff Infrastructure Engineer Jobs
3,518 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current staff infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are looking for a Staff Full Stack Engineer to build the user experience for the Okta Recovery Vault (ORV) . The ORV is a critical new enterprise-grade capability designed to protect against the accidental or malicious deletion of identity objects. This role is essential to our "Land big, grow fast" strategy, by providing customers with a reliable safety net for their data. You will be responsible for designing and implementing a highly intuitive self-service recovery UI . This interface must empower admins to perform complex operations—such as bulk restorations of hundreds of users, managing configurable retention policies, and navigating detailed post-restore checklists—with confidence and ease. Job Duties and Responsibilities UI Development: Lead the implementation of the Recovery Vault UI, ensuring a seamless experience for authorized admins via the Okta Admin Console. Complex Workflows: Build interactive interfaces for granular and bulk recovery operations, including async job progress indicators and success notifications.<
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Technical Recruiting team partners with Engineering, Security, Data, and other technical orgs to build the teams behind Robinhood’s products and infrastructure. We work closely with hiring leaders to plan hiring, build strong pipelines, and deliver a high-quality, consistent candidate experience. We care about clear communication, strong partnerships, and continuously improving how we hire. As a Technical Recruiter, you’ll own full-cycle recruiting across a range of technical roles. You’ll partner with hiring managers and senior leaders in Engineering to define what great looks like, develop effective sourcing strategies, and drive consistent pipeline activity. You’ll collaborate closely with peers to share insights, improve workflows, and keep hiring moving efficiently. Th
About the Team ChatGPT is a rapidly evolving system: new capabilities ship continuously, product surfaces change quickly, and usage patterns shift week-to-week. Supporting that pace requires infrastructure that can handle real production constraints—high concurrency, unpredictable traffic patterns, complex dependency graphs, and frequent change. The ChatGPT Infrastructure team builds and operates the platforms that enable fast iteration without compromising performance or reliability. We design shared systems, data paths, rollout mechanisms, and reliability guardrails that teams rely on to ship changes to ChatGPT at scale. We focus on high-leverage infrastructure: primitives and “golden paths” that incorporate operational lessons as defaults, so engineers don’t need to rediscover failure modes, latency pitfalls, or integration issues each time they build something new. About the Role We’re hiring Senior and Staff Engineers to design and build infrastructure systems that underlie ChatGPT and multiply the effectiveness of teams building user experiences. This is not a support-only role. It’s a platform-building role: you’ll define interfaces, develop core abstractions, and create tooling to make safe, fast iteration the norm. Your work will reduce friction, prevent regressions, improve performance, and ensure systems scale gracefully as the product grows. Where You Can Have Impact You might work on one or more of the following areas (without being restricted to any single area): Platform foundations & frameworks: Core libraries, service frameworks, and shared components that standardize system building, integration, and evolution. Scalability & performance primitives: Patterns and infrastructure that reduce tail latency, improve throughput, and keep costs predictable as demand increases. Reliability guardrails: Mechanisms that prevent outages by design—rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks, and robust regression contr
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Developer Infrastructure org is the engine behind Robinhood's entire engineering organization — a collection of tightly integrated teams whose collective mission is to make every engineer at Robinhood faster, more reliable, and exponentially more productive. The org spans four major teams: DevX (Developer Experience), TestX (Test Infrastructure), Backend Platform, and Mobile Platform. DevX owns Robinhood's monorepo and Bazel-based build infrastructure — the critical layer between a developer writing code and that code being ready to ship — along with the company's full CI/CD pipeline and remote build execution cluster. TestX owns the infrastructure behind Robinhood's entire test experience: the integration test environments, and personal development environments that serve as miniature simulations of the full Robinhood system, giving engineers a safe, isolated space to test their code end-to-end before it ever touches production. Backend Platform and Mobile Platform own the core language runtimes, libraries, IDEs, and developer toolchains across Python, Go, TypeScript, Swift, and Android. Together, these teams share a single north star: leveraging AI and agentic systems
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Data Product Platform is mission-critical to accelerating data-driven decision-making at Pinterest on the foundation of 100s of thousands of tables and an exadata-scale data warehouse. We strive to provide effortless, efficient, and reliable data products and platforms that power the entire company. We achieve this by investing in three core areas: Data Warehouse: Building and managing the foundational data warehouses that enable key analyses across both our core engagement and monetization products. Analytical Velocity: Creating powerful analytical tools that empower internal data users to leverage our vast data assets and capable infrastructure effectively. Data Governance: Defining and implementing the data governance policies and tools necessary to ensure the responsible, efficient and compliant storage and handling of all data. We are s
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Contribute in and provide strong support for model training pipelines, ship state of the art models to production, and bridge the gap between research and production. We have one of the highest ratio of compute to engineers in the world. We do not delineate strongly between engineering and research. Everyone will contribute to writing production code and supporting our research effort depending on individual interest and organizational needs. We have all the compute, data, and talent available for you to do your best work. Please Note: We have offices in London, Paris, Toronto, San Francisco and New York but also embrace being remote-friendly! There are no restrictions on where you can be located for this role. As a Member of Technical Staff, you will: Design and write high-performant and scalable software for training. Improve our training setup from an infrastructure and codebase performance standpoint. Craft and implement tools to speed up our training cycles and improve the overall efficacy of our training infrastructure Research, implement, and experiment with ideas on our supercompute and data infrastructure
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by leading the design of high-performance, scalable and reliable machine learning systems? Do you want to set technical direction and help shape the next generation of AI platforms powering advanced NLP applications? We are looking for a Lead Member of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will serve as a key point of contact for customers, leading the design of customized deployments to meet their specific needs, and mentoring engineers to raise the technical bar across the team. You may be a good fit if you have: 8+ years of engineering experience running production infrastructure at a large scale, with a track record of technical leadership Demonstrated experience leading the architecture
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us—that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale—from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role Nuro's autonomous vehicle platform has to work - not just in the easy cases, but in the hard ones. The ones we design for, simulate at scale, and deliberately break to understand. We're looking for a Senior/Staff Systems Engineering TPM who owns the technical substance of how we validate autonomy: what we test, why we test it, and whether our coverage actually means something. This is not a planning or scheduling role. You'll work embedded with Autonomy, Simulation, and Systems Engineering teams to drive the technical rigor behind our validation program — defining scene sets, structuring fault injection campaigns, and ensuring our simulation coverage is meaningful, traceable, and systematically growing. The program infrastructure is owned elsewhere; your job is to make sure what's inside it is technically sound. About the Work Scenario & Coverage Strategy Understanding the taxonomy of challenging scene sets used for auto
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr. Staff Network Engineer to join our team. This is a fully remote capacity in the Netherlands role, reporting to the Manager, Network Engineering in the Cloud Ops - Network Engineering department. This position involves a mix of infrastructure, project management, and network engineering responsibilities. You will join our global team as an integral member, taking ownership of the deployment, monitoring, and ongoing operation of our worldwide production infrastructure across diverse data center locations. We are seeking a high-trust collaborator with a background in large-scale enterprise or telecom networks who approaches challenges with a growth mindset and a commitment to achieving engineering excellence. This role requires active participation in our technical on-call rotations, which include support during weekends and holidays to ensure continuous service reliability. What you’ll d
Get new staff infrastructure engineer jobs by email
Daily job updates · Unsubscribe anytime