Jobiba hiring network

Staff Production Engineer Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current staff production engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Z
Zscaler
📍 CaliforniaFull-timeRemoteFrom $152K/yr
3 days ago

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr. Production Engineer to join our team. This role is available as a hybrid opportunity 3 days a week in San Jose, CA or Remote reporting to Production Engineering in the Cloud Infrastructure & Operations department. Join Zscaler to be a force multiplier for the reliability of a global platform processing 200+ billion transactions daily across tens of millions of enterprise users. In this role, you will provide the technical vision and hands-on execution to drive an "automation-first" culture across the company. By maturing our observability and architectural standards, you will directly reduce our Mean Time to Mitigate (MTTM) and shape the scalability of our globally distributed, multi-cloud infrastructure. What you’ll do (Role Expectations) Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments Drive an "automation-first" culture by wr

REMOTEpythonawsazure
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking a highly technical Lead Release Engineer with a strong engineering foundation to lead the end-to-end release lifecycle of our storage products. You will serve as the bridge between Development, QA, and Product Management — ensuring that complex storage stacks are delivered with high quality and predictable cadences. Unlike traditional project-based release management, this role demands deep hands-on expertise across CI/CD orchestration, codeline management, system-level triaging, fleet operations, and the engineering rigor required for data-critical products. You will own the health of our release pipelines, lead triage war rooms, drive automation initiatives, and participate in on-call rotations to keep CI and test-orchestration infrastructure running reliably. You will also build developer-facing tooling, manage HW test fleet operations, and maintain high-quality integration workflows across our code lines. WHAT YOU'LL DO Release Orchestration: Own the end-to-end release process for storage software and firmware, from development to GA (General Availability). CI/CD Leadership: Design and build optimized pipelines and tools to scale code management and merge operations. Work closely with systems such as Jenkins, test frameworks, Premerge, Orchestrator, and related developer productivity tooling to keep the codeline healthy and actionable. Technical Triaging: Act as the primary technical poin

pythonawsci/cd
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Overview We are looking for a strong software and platform engineer to join our Production Engineering team in Bangalore as an individual contributor in FA ProductionEng APJ. This role will help build and operate internal platforms that improve how we provision, observe, govern, and troubleshoot engineering infrastructure at scale. The fleet management use cases that give teams a single place to understand and operate the test infrastructure. If you enjoy building internal platforms that remove friction, improve visibility, and make engineering teams faster and more effective, this role is for you. Why This Role Is Unique This is not a typical application development role.You will work on internal platforms that directly shape how engineering teams consume and manage shared infrastructure. The role spans platform engineering, workflow automation, observability, API-driven services, and infrastructure lifecycle management. The right candidate will work on systems such as: Self-serviceability workflows and lease-based testbed governance. Developer Platform dashboards and APIs used for triage, visibility, and product trend observation. Testbed and workflow orchestration across fleet management domains. Impact This role is a high-leverage engineering investment. The work will improve how engineering teams provision testbeds, understand failures, operate shared infrastructure, and move faster with less friction. Better

awskuberneteslinux
View job →
G
17 days ago

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... GEDE (Global Edge and Domains Engineering) keeps GoDaddy's domains, edge, and aftermarket platforms running for millions of customers. Within GEDE, the Reliability Engineering team (Domains Production Engineering) is the group that gets the call when something breaks — and, more importantly, builds the systems that mean it breaks less often. We own the monitoring, compliance, and patching toolchain for the whole Domains infrastructure footprint, provide advanced incident support, and are actively modernizing how the org detects, diagnoses, and even auto-remediates issues with AI-assisted tooling! What you'll get to do... Build and evolve observability using Prometheus/Mimir, the Grafana LGTM stack, Elastic/OTEL, and Site24x7 — closing gaps across the org. Steer our cloud migration journey and bolster our efforts to keep the services reliable and performant. Own patching compliance and vulnerability remediation at scale across a mixed on-prem + AWS fleet, hitting hard SLA targets. Operate and extend our multi-tenant Kubernetes/ArgoCD platform, including the migration of core services. Contribute to our AI/automation initiatives: auto-generating runbooks from Prometheus alerts, ServiceNow change-risk scoring, and other tooling that reduces toil for the whole team Consult with partner dev teams on metrics, alert thresholds, and monitoring standards — this is a platform-enablement role, not just a ticket queue. Mentor other engineers on the team and help mature our operational practices. Your experience should include

pythonsqlpostgresql
View job →
B
Biohub
📍 Redwood CityFull-timeHybrid$241K – $331K/yr
1mo ago

Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping peop

pythonkubernetesgit
View job →
J
Jumio
📍 IndiaFull-timeRemote
3 days ago

Role Purpose We’re looking for a Staff/Senior Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Experience: 5+ years of industry experience in Machine Learning, with at least 3 years dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native: Hands-on ex

REMOTEpythonawsrest
View job →
J
3 days ago

Machine Learning Engineer IV – (Computer Vision) We’re looking for a Staff/Senior Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Strong industry experience in Machine Learning, dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native: Hands-on exper

REMOTEpythonawsrest
View job →
M
Mongodb
📍 TorontoFull-timeFrom C$144K/yr
1mo ago

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based

pythonmongodbaws
View job →
M
1mo ago

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based

pythonmongodbaws
View job →
R
Roblox
📍 San MateoFull-timeFrom $326.1K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when

pythonjavaaws
View job →
P
Pendo
📍 New YorkFull-time$300K – $325K/yr
1mo ago

The Team + The Role Our Emerging Team is focused on building AI Products for our product experience (PX) platform. We build from the ground up to explore, prototype, and ship AI-native experiences that change how software teams understand and serve their users. This is not an AI layer added to existing product; it is a deliberate bet on what product intelligence looks like next. The team operates with high autonomy, moves quickly, and builds products without clear precedents. As a Staff Software Engineer (AI), you will sit at the intersection of deep technical capability and strong product judgment. You will design and build production-grade AI systems, including RAG pipelines, agentic workflows, and LLM-powered features, while making clear tradeoffs across prompting, fine-tuning, architecture, evaluation, and deployment. You will also partner closely with product, design, and engineering stakeholders to frame the right problems and communicate technical decisions clearly. This role is based in our New York office. What this looks like day-to-day Applied AI systems: Design and build AI-native systems, including RAG pipelines, agentic workflows, and LLM-powered product features. You will take ideas from prototype through production and ensure they can support real users. Model strategy: Make principled decisions about when to prompt, when to fine-tune, and when to use a different technical approach entirely. You will explain those tradeoffs clearly to engineers and non-engineers. Evaluation and guardrails: Instrument and evaluate model outputs rigorously by defining evaluation frameworks and identifying hallucinations early. You will implement guardrails that hold up under real-world usage and load. Productionize AI ownership: Own model deployment, monitoring, latency optimization, cost management, and reliability at scale. You will ensure AI systems are observable, performant, and production-ready. Full-stack delivery: Contribute across the stack when needed to get

C
Clickup
📍 United StatesFull-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer – AI Product to join our ClickUp Engineering team. In this role, you will drive the development of intelligent, user-facing features that leverage the latest advancements in AI and large language models (LLMs). You will work closely with product, design, and engineering teams to deliver seamless, impactful AI-powered experiences that delight our users and differentiate ClickUp in the productivity space. This is a product-focused engineering role requiring deep expertise in AI, LLMs, and building scalable, production-ready applications. Key Responsibilities: Lead the design, development, and deployment of AI-powered features and products that directly impact ClickUp users. Collaborate with product managers, designers, and engineers to identify opportunities for AI-driven innovation and translate user needs into technical solutions. Integrate and orchestrate multiple LLMs and AI models to deliver robust, context-aware, and personalized user experiences. Prototype, test, and iterate on new AI features, leveraging user feedback and data to drive continuous improvement. Ensure the reliability, scalability, and performance of AI-powered features in production environments. Stay at the forefront of AI research and product trends, incorporating the latest advancements into ClickUp’s product roadmap. Address AI privacy, security, and compliance challenges, ensuring responsible and ethical use of AI in user-facing applications. Mentor and guide other engineers in best practices for building AI-powered products. Qualifications: Proven experience bui

javascripttypescriptpython
View job →
O
Okta
📍 TorontoFull-timeFrom C$168K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team : Have you ever considered what powers the intelligent features behind seamless product experiences? The GenAI team is at the forefront of enabling AI-powered security and intelligent innovation across our organization. From crafting AI powered security services, to intuitive generative AI-powered chat experiences that provide instant product support, and developing the best developer experience around authentication for generative AI and AI agents, our team is instrumental in bringing the transformative power of AI to life. We collaborate closely with the Machine Learning team and various product teams to ensure the seamless and secure delivery of AI-enhanced features that provide real value to our users. The Opportunity : As a Staff Machine Learning Engineer on the Generative AI team, you will help shape, architect, and accelerate our Generative AI strategy by contributing across the stack of model development, infrastructure, and platform services. You’ll drive design and implementation of production-ready AI/ML systems at scale: ranging from LLM-powered features to reusable components that other teams across Okta can build on. You will have the opportunity to: Architect, design, and deploy robust Machine Learning & GenAI systems, ensuring seamless integration with diverse platform services and establishing scalable LLMOps pipelines in production. Drive technical decision making while striving to hit the right balance between factors s

typescriptpythonaws
View job →
A
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Host Pricing & Settings team builds the platform and tools that help hosts run their business — with pricing strategies informed by market intelligence, comparable listings, and demand signals. We partner with Search, Listings, Tax, and Payments to ensure our guidance is accurate, timely, and trusted. Behind every pricing recommendation is a sophisticated ML system undergoing a fundamental rearchitecture. Our north star: a serving infrastructure where training, inference, and evaluation are consistent by design — features from a centralized store, model composition in one place, and backfills available on demand so data scientists and MLEs can evaluate candidates in days, not weeks. The Difference You Will Make: As a senior technical individual contributor, you will own the technical strategy for the full Modeling → ML Serving → API interface across the Host Pricing org. Although you will be at one of our highest levels of seniority, all individual contributors at Airbnb are Software Engineers — you are expected to be hands-on and contribute code. Define the architecture and contracts governing how models move from development to production — feature store design, model schema management, online/offline inference consistency, and multi-version support. Lead the buildout of a unified serving stack that eliminates per-model one-off implementations and gives data scientists a turnkey path from training to production. Architect backfill and evaluation infrastructure so the modeling team can simulate production inference over historical data in days, not weeks. Establish do

REMOTEpythonjavaai
View job →
C
Clickup
📍 United StatesFull-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay current with ad

typescriptpythonaws
View job →
🔔

Get new staff production engineer jobs by email

Daily job updates · Unsubscribe anytime