The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim
Jobiba hiring network
Distributed Systems Engineer Data Platform Delivery Database Retrieval Jobs
1,301 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer data platform delivery database retrieval jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We are seeking an experienced Quantitative Developer to join our Markets Quantitative Analytics team, partnering closely with Quantitative Analysts, Traders, and Technology professionals to build the next generation of pricing, risk, and analytics platforms. This is a hands-on technical role for a highly skilled software engineer with a passion for quantitative finance. You will be responsible for designing and delivering high-performance, scalable solutions that support front office trading businesses across asset classes. The role offers the opportunity to work on complex quantitative challenges, modern engineering practices, and large-scale distributed systems while helping shape the strategic direction of Citi's quantitative technology platform. Successful candidates will combine strong software engineering expertise with an understanding of quantitative methodologies and financial markets, translating sophisticated mathematical models into robust, production-grade solutions. Key Responsibilities Design, develop, and maintain high-performance pricing, risk, and analytics libraries used across Global Markets. Partner with Quantitative Analysts to transform research models and prototypes into scalable, production-quality software. Build and optimize quantitative applications using modern C++ and Python, applying strong software architecture and engineering principles. Own the full software development lifecycle, including requirements gathering, design, implementation, testing, deployment, and ongoing support. Drive engineering excellence through CI/CD adoption, automated testing, code reviews, and software quality best practices. Develop and maintain market data platforms and data pipelines supporting analytics, pricing, and risk workflows. Work with infrastructure teams to leverage distributed computing, cloud technologies, and scalable arc
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the foundational infrastructure that will define how the world thinks about and deploys AI, and we want the sharpest, most curious people to help us do it. As a member of our Analytics & Data Insights team, you'll tackle the kind of problems that don't have textbook answers yet, launch products that didn't exist a year ago, and help enterprises understand what foundational AI actually means for their bottom line. As a Data Engineer, you will: Work directly on new customer experiences built on one of the most advanced AI systems in the world Collaborate daily with researchers and engineers who are some of the best in the world at what they do Run implementations end-to-end and see initiatives through to real outcomes Partner across research, marketing, sales, and finance to help define how Cohere grows, with your recommendations feeding directly into products and strategy You may be a good fit if you have: 5+ years of experience working on production-grade data processing systems Strong command of Python and SQL Experience with distributed data processing frameworks such as Apache Beam, Spark, or
Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been powering MLEs, researchers, data scientists and operators for fast and automatic training and evaluation of LLM's, as well as evaluation of data quality. Scale is uniquely positioned at the heart of the field of AI as an indispensable provider of training and evaluation data and end-to-end solutions for the ML lifecycle. You will work closely across Scale’s ML teams and researchers to build the foundation platform that supports all our ML research and development. You will be building and optimizing the platform to enable our next generation of LLM training, inference and data curation. If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you! You will: Build, profile and optimize our training and inference framework Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation Research and integrate state-of-the-art technologies to optimize our ML system Ideally you’d have: Strong excitement about system optimization Experience with multi-node LLM training and inference Experience with developing large-scale distributed ML systems Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc. Strong written and verbal communication skills and the ability to operate in a cross functional team environment Nice to haves: Demonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the positi
When 5% of Indian households shop with us, it’s important to build data-backed, resilient systems to manage millions of orders every day. We’ve done this – with zero downtime! 😎 Sounds impossible? Well, that’s the kind of Engineering muscle that has helped Meesho become the e-commerce giant that it is today. We value speed over perfection, and see failures as opportunities to become better. We’ve taken steps to inculcate a strong ‘Founder’s Mindset’ across our engineering teams, making us grow and move fast. We place special emphasis on the continuous growth of each team member - and we do this with regular 1-1s and open communication. Tech Culture We have a unique tech culture where engineers are seen as problem solvers. The engineering org is divided into multiple pods and each pod is aligned to a particular business theme. It is a culture driven by logical debates & arguments rather than authority. At Meesho, you get to solve hard technical problems at scale as well as have a significant impact on the lives of millions of entrepreneurs. You are expected to contribute to the Solutioning of product problems as well as challenge existing solutions. Meesho’s user base has grown 4x in the last 1 year and we have more than 50 million downloads of our app. Here are a few projects we have completed last year to scale oursystems for this growth: ● We have developed API gateway aggregators using frameworks like Hystrix and spring-cloud-gateway for circuit breaking and parallel processing. ● Our serving microservices handle more than 15K RPS on normal days and during saledays this can go to 30K RPS. Being a consumer app, these systems have SLAs of ~10ms ● Our distributed scheduler tracks more than 50 million shipments periodically fromdifferent partners and does async processing involving RDBMS. ● We use an in-house video streaming platform to support a wide variety of devices and networks.
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Storage Services team operates Pinterest's scalable online structured data storage platform—supporting both SQL-based table models and complex graph data structures—managing 100TB+ datasets and serving over 1.5M queries per second across many of Pinterest's most important products. We're looking for an exceptional Staff Software Engineer to lead the technical strategy and execution of our storage infrastructure initiatives, defining how these systems are designed, built, and operated. You'll drive innovation across distributed SQL, high-throughput/low-latency query processing, graph workloads, and the developer experience for storage clients. What you’ll do: Provide technical guidance and direction to a high-performing team building reliable, performant, and cost-efficient storage systems that operate at massive scale and power busines
The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Training & Serving team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: distributed training of foundation models, serving at scale, designing the user experience. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Training & Serving team, directly managing 10+ engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage, infrastructure and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: Previous experience (1+ years) leading software engineering teams, as a tech lead or people manager Strong technician with a mix of backend, data engineer and infrastructure experience who is interested in remaining a hands-on leader Excellent leader with strong
The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the data infrastructure behind some of the most demanding AI training workloads in the world, and we want sharp, curious people to help us do it. In this role, you'll build and maintain the high-performance data layer our Modeling teams rely on for training and evaluation jobs. As a Software Engineer, Data Infrastructure, you will: Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it. Collaborate daily with researchers and engineers who are some of the best in the world at what they do. You may be a good fit if you have: 4+ years of experience working on data storage infrastructure Strong command of Python Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.) The ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink [Nice-to-have] Familiarity with modern analytics tooling such as BigQuery, Airflow, or dbt Genuine excitement about AI.
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale. About the Work Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration. Design and operate large-scale data pipelines - batch and strea
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale. About the Work Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration. Design and operate large-scale data pipelines - batch and strea
About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role is hybrid requiring 2 days in office at our San Francisco or NYC hub every Tuesday & Wednesday. About the Role Machine Learning is a cornerstone at Taskrabbit, and we're looking for a Staff Machine Learning Engineer to join our team and lead the next phase of our customer retention strategy. This is a critical, full-stack role for an individual who is passionate about the end-to-end lifecycle: from initial research and model development to building the robust systems that power repeat customer engagement and lifetime value growth at scale. Taskrabbit's greatest growth opportunity lies in deepening customer relationships and accelerating repeat purchases. Our most valuable customers are those who return frequently, discover new service categories, and increase their spending over time. There's significant untapped potential in the marketplace: repeat customers spend 3-5x more than one-time users, and category expan
About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role is hybrid requiring 2 days in office at our San Francisco or NYC hub every Tuesday & Wednesday. About the Role Machine Learning is a cornerstone at Taskrabbit, and we're looking for a Staff Machine Learning Engineer to join our team and lead the next phase of our customer retention strategy. This is a critical, full-stack role for an individual who is passionate about the end-to-end lifecycle: from initial research and model development to building the robust systems that power repeat customer engagement and lifetime value growth at scale. Taskrabbit's greatest growth opportunity lies in deepening customer relationships and accelerating repeat purchases. Our most valuable customers are those who return frequently, discover new service categories, and increase their spending over time. There's significant untapped potential in the marketplace: repeat customers spend 3-5x more than one-time users, and category expan
About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! Prior to applying please note: W e are currently unable to provide visa sponsorship for this position (including H-1B, OPT, or other employment-based visas). Candidates must be legally authorized to work in the United States without employer sponsorship now or in the future. This role is hybrid requiring 2 days in office at our San Francisco hub every Tuesday & Wednesday (located at 130 Sutter St). About the Role Machine Learning is a cornerstone at Taskrabbit, and we’re looking for a Staff Machine Learning Engineer to take technical ownership of our core ranking system. Every job request on the platform flows through it, making this one of the most consequential ML systems we run. This is a hands-on technical leadership role. You’ll operate as the primary architect and engineer for the ranking system — defining the system direction, driving the roadmap, solving the hardest problems, and creating leverage for the engi
Get new distributed systems engineer data platform delivery database retrieval jobs by email
Daily job updates · Unsubscribe anytime