About the Team The Enterprise Identity team builds the identity foundation that enables organizations to adopt and use OpenAI products securely and reliably. The team owns the enterprise identity stack, including SSO, SCIM, tenant architecture, and identity capabilities across the enterprise admin experience and OpenAI's growing multi-product portfolio. About the Role We are looking for a hands-on senior technical leader to own the architecture and evolution of OpenAI's Enterprise Identity systems. You will set the long-term technical vision for the entire stack, establish shared identity primitives across products, and be accountable for systems that are foundational to our enterprise business. This role requires operating well beyond a single service or feature area. You will identify the most consequential architectural investments, align teams around durable solutions, and ensure our identity platform meets an exceptionally high bar for scale, availability, latency, and security. This role will be based in our San Francisco or Mountain View office. In this role, you will: Own the technical vision and architecture for the Enterprise Identity stack, including SSO, SCIM, tenant architecture, groups, permissions, and identity capabilities in enterprise administration surfaces. Lead the design and evolution of highly available, latency-sensitive identity systems serving a large and diverse global enterprise customer base. Establish common identity models and primitives that work consistently across OpenAI's products and enable the organization to scale. Set a high security bar by anticipating abuse cases, failure modes, and the long-term implications of new capabilities. Drive alignment across enterprise product, infrastructure, and security partners, resolving ambiguity and influencing roadmaps beyond the immediate team. Provide technical leadership to senior engineers and raise the quality of architecture and execution across the broader organization. You might thr
Jobiba hiring network
Staff Software Reliability Engineer Data Platform Jobs
3,518 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current staff software reliability engineer data platform jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team Join the engineering teams that bring OpenAI’s ideas safely to the world! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We’re seeking Software Engineers who can solve complex, high-impact problems across our stack. In this role, you’ll join a nimble team driving the deployment of OpenAI’s technology into new environments and infrastructure that power critical missions in the public sector. You’ll work cross-functionally with product, security, and compliance teams to build the functionality needed to deliver a scalable, reliable platform. You’ll also partner directly with customers to design and build new products and features that create real-world impact. From launching net-new capabilities to optimizing how we serve inference in unique, high-stakes environments, this role offers both breadth and technical depth—giving you the opportunity to shape the future of OpenAI’s technology where it matters most. This role is based in Washington D.C., San Francisco, CA or Seattle, WA. Occasional travel to customer sites is required for this role. In this role, you will: Own the development of new customer-facing ChatGPT and OpenAI API features end-to-end, both on-premises and in the cloud, for our public sector customers. Partner and directly embed with teams across the business, including engineering, security, and compliance, to enable our products to work within the unique constraints of new environments. Talk to users to understand their problems and design solutions to address them Work with the research team to get relevant feedback and iterate on their latest models, developing solutions specific for public sector customers at both the model & data
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build the future of data. Join the Snowflake team. The Snowflake Machine Learning Platform team’s mission is to enable customers to bring their machine learning and deep learning workloads to Snowflake. Our customers want to build powerful models with the ever-increasing data in Snowflake but face several challenges including infrastructure optimizations, orchestration, performance, and security. The team aims to solve these challenges by building highly integrated platform solutions that are simple, secure, and enable end-to-end ML workflows. We are on an early journey to build the most scalable machine learning and data platform without sacrificing the benefits of a single platform and governance. We are looking for outstanding technical leaders who will join our ML Platform team to build the next-generation platform and play a pivotal role in this journey by understanding Snowflake’s core platform architecture and evolving it to enable state-of-the-art machine learning and LLM workloads. Join us to define strategies, set technical directions, design and execute, engage and deliver innovation, and unlock the power of AI for thousands of enterprise customers. This position is based in Menlo Park, CA, and Bellevue, WA. RESPONSIBILITIES : Help define and own the roadmap, wor
About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! In September 2026 we raised a $350 million Series E at a $3.5 billion valuation , and we are scaling our engineering and research teams to meet demand. The role Frontier AI data is expensive to make and hard to measure. Every task we deliver is tested against the strongest models, often through many long-running agent rollouts. Your job is to make that process faster, cheaper, and more rigorous with ML and AI You will be one of the early members of ML & Research Engineering at Snorkel. You will study how frontier-grade data is generated and evaluated, form hypotheses, validate them against real production data, and ship the winners at scale. You will shape the discipline's direction, its standards, and the team that grows around it. What you'll work on Efficient agentic evals. Cut the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating. AI model routing. Route every eval and judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution. Fine-tuned small models. Fine-tune and serve open-weight models (LoRA and other
About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr. Staff Software Development Engineer-AI Security to join our team. This is a Hybrid (based in San Jose, CA or Bellevue, WA with a 3 days in office requirement) role, reporting to the Director of Software Engineering in the Emerging Tech department. You will be responsible for designing and implementing core infrastructure components and distributed systems, serving as a foundational architect for our AI security solution. This high-impact role focuses on scaling security infrastructure to support hundreds of millions of users, collaborating with stakeholders across the development lifecycle to drive innovation and technical excellence. What you’ll do (Role Expectations) Architect, develop, and optimize a low-latency, high-throughput AI Security plane utilizing Rust, specifically leveraging its async/await model for highly efficient I/O and service-oriented architecture Build resi
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Our Merchants organization focuses on onboarding and activation, catalog quality and health, merchant presence (including creator/UGC and trust signals), and ads upsell—helping merchants build healthy catalogs, reach the right audiences, and unlock clear paths to growth. In this role, you’ll lead high-impact engineering at the intersection of commerce platforms, catalog systems, and AI-native experiences. This includes applying ML, AI-assisted workflows, and GenAI where appropriate to improve metadata quality, merchant tooling, and operational efficiency. You’ll partner closely with Engineering Managers, Product Managers, Data Scientists, and other senior engineers to deliver systems with measurable business and user impact. What you’ll do: Own end-to-end technical delivery for cross-team initiatives—from problem framing and technical stra
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . What we’re looking for: Millions of people across the world come to Pinterest to find new ideas every day. It’s where they get inspiration, dream about new possibilities and plan for what matters most. Our mission is to help those people find their inspiration and create a life they love. In your role, you’ll be challenged to take on work that upholds this mission and pushes Pinterest forward. You’ll grow as a person and leader in your field, all the while helping Pinners make their lives better in the positive corner of the internet. We are looking for a passionate, inquisitive, and well-rounded Sr. Staff Backend Engineer to join the Pinterest Assistant team. The team is building a visual-first, AI-powered companion that helps Pinners go from inspiration to action across shopping, search, and discovery. As the technical leader for our backend p
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest helps more than 600 million people discover new ideas to help design their life. Our users come to Pinterest to explore ideas and run more than 80 billion search queries every month. Many of these queries represent exploratory search and shopping intent and are broad, which means that the search system should be able to deeply understand this intent, then help people explore content, personalize results, harness visual and multimodal signals effectively, and show the most engaging content up front. All of this means that Pinterest Search presents a unique challenge quite unlike other search systems and the opportunity to innovate on a product that only Pinterest can build. We are looking for an exceptional Senior Staff Engineer to lead Search’s product and front-end strategy across iOS and other platforms such as Android and Web, spann
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Data Product Platform is mission-critical to accelerating data-driven decision-making at Pinterest on the foundation of 100s of thousands of tables and an exadata-scale data warehouse. We strive to provide effortless, efficient, and reliable data products and platforms that power the entire company. We achieve this by investing in three core areas: Data Warehouse: Building and managing the foundational data warehouses that enable key analyses across both our core engagement and monetization products. Analytical Velocity: Creating powerful analytical tools that empower internal data users to leverage our vast data assets and capable infrastructure effectively. Data Governance: Defining and implementing the data governance policies and tools necessary to ensure the responsible, efficient and compliant storage and handling of all data. We are s
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Ads Platform team builds the AI-powered growth engine behind Airbnb's paid marketing across Google, Meta, and the broader advertising ecosystem. We build the platform and intelligent systems that power paid campaigns — from creative to optimization to audience — at scale across every channel, with a strong commitment to brand safety, quality, and human oversight. Our ambition is bold: to build increasingly autonomous, AI-driven marketing systems that grow Airbnb's business. We sit at the intersection of engineering, marketing, and AI. You'll partner closely with Marketing, Product, and Design, work alongside Airbnb's ML/AI and data platform teams, and integrate deeply with the core product to turn a rich understanding of the customer journey into differentiated, high-performing paid experiences. The Difference You Will Make: As a Staff Software Engineer on the Ads Platform team, you will lead the technical strategy and hands-on execution of Airbnb's next-generation, AI-powered Growth Engine for paid channels (Google, Meta, and beyond). You will architect systems that autonomously manage the full lifecycle of paid campaigns — creative generation and rotation, bidding and budget optimization, audience management, and feed intelligence — at scale. Success looks like measurable gains in marketing efficiency (ROAS/CAC) and campaign velocity as we advance the platform along its AI-assisted → agentic → autonomous maturity curve, with the quality, brand safety, and human-in-the-loop controls that let us operate autonomously with confidence. A Typical Day: Set technical direction f
MongoDB is seeking a Staff Software Engineer to join the Atlas Clusters Organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. This engineer will also work on our Atlas Data Federation & Archiving product. Atlas Data Federation & Archiving allows customers to move data from hot to cold storage and run federated queries over that data. We are forming a new Atlas Clusters team in the Dublin area. We are looking to speak to candidates who are based in Dublin and would like a hybrid or in-office working model. What you’ll do Build and design new features for MongoDB Atlas and Atlas Data Federation & Archiving Contribute to and lead complex technical projects Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 10+ years of professional software development experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Go, Java, C#, etc) Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has led the launch of a new module and maintained it in production Is eager to solve tough problems Has excellent communication skills Is curious, collaborative, and motivated Success Measures In 3 months, you'll have shipped code into production and c
We are looking for an experienced Staff Software Engineer with a strong background in building software frameworks, LLM eval frameworks and team leadership. You will lead a talented team building a complex product suite using a modern stack. The ideal candidate is a hands-on technical leader who can drive architectural decisions, mentor engineers, and collaborate closely with product management to deliver a product that solves our customers' most challenging modernization validations problems, ensuring that new applications have the same functionality and performance as their legacy counterparts. This role will be based in our India office in Gurgaon and offers a hybrid working model. The ideal candidate for this role will have 8+ years of commercial software development experience with at least one JVM language such as Java, preferably using the Spring ecosystem and paired with strong Python experience 2+ years of experience leading, coaching, and mentoring a team of software engineers to achieve high-impact results Solid experience in software architecture and development Familiarity with evaluation frameworks (such as DeepEval, Ragas, and Promptfoo) and harnesses, alongside benchmark datasets, scoring/grading approaches, and LLM-as-a-judge methodologies Experience in application/database modernization with proven ability to design and implement frameworks that ensure database state equivalence across different implementations Must understand change data capture, event-driven database interception (MongoDB listeners, RDBMS triggers), and state comparison algorithms with pattern-based exclusions Extensive experience with relational and document data modeling and hands-on experience with at least one SQL database (Postgres, MySQL, etc) and at least one document database (e.g. MongoDB) Good understanding of algorithms, data structures and their time and space complexity Curiosity, a positive attitude, and a drive to continue learning Excellent verbal and written comm
We’re looking for a Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical leader within the APM organization, driving GenAI/machine learning projects from concept to production. Build and benchmark GenAI/ML models using state-of-the-art techniques. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, scaling challenges, and evolving requirements with clear technical direction. Actively mentor engineers and influence engineering culture through leadership in design reviews, technical talks, and working groups. Who You Are: You have a BS/MS/PhD in a scientific field or equiva
The worldwide data management software market is massive (IDC forecasts it to be $138 billion by 2026). At MongoDB, we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the center of innovation and creativity. MongoDB is seeking a Sr. Staff Software Engineer to join the Atlas Core Data Services organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product, along with the API Platform and Developer Tools. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Core Data Services organization builds the software that manages the Atlas cluster infrastructure hosted on the three major cloud providers (AWS, Azure, and GCP), as well as the software that manages the MongoDB database hosted on that infrastructure. We are constantly challenged to design features that ensure Atlas clusters are secure, available, durable, and performant while running large-scale, critical workloads. The Sr. Staff Engineer in this role will drive innovation across the organization and the company, setting technical standards and direction that enable future growth and velocity. We are looking for engineers with the experience and high standards needed to lead at that scale. Our organization champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies systems expertise to build the foundational infrastructure of a popular database, join us. Let's build a faster, more reliable, and highly scalable database platform together. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Define standards and vision for the mission-critical Atlas SaaS data
Get new staff software reliability engineer data platform jobs by email
Daily job updates · Unsubscribe anytime