At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization . Our mission is to build the next generation of high-performance and intelligent inference systems . We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads. Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost. Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering , using AI not only as the workload we optimize, but also as a tool to accelerate system deve
Jobiba hiring network
Distributed Systems Engineer Jobs
1,301 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations Atlanta, US Austin, US Denver, US New York, US Toronto, Canada Washington DC, US Seattle, WA Remote candidates within North America will also be considered. About the Role The Security Platform team is an infrastructure/developer tools group tasked with building and operating powerful, resilient, and secure infrastructure and systems that enable other engineering teams to deliver products to our customers efficiently and secure
We deliver foundational systems that shape the future of how technology is used in a top performing quantitative equity fund that manages over $78+ billion USD in financial assets. This is a fantastic opportunity in the exciting intersection of finance and technology where investment decisions are made using technology. Quantitative equity funds use programmed investment strategies and as a result, our technology team is crucial to its success. The team is headquartered and deeply rooted in West Coast Vancouver. We place high value on maintaining an entrepreneurial spirit and creating a culture where each of us has opportunities to succeed. What You Will Do The technology infrastructure team plays an essential role through innovative technologies on our hybrid (on-premise and cloud based) platform: distributed computing, petabyte-scale data storage, containerization, non-traditional high-performance databases, process orchestration, monitoring, data visualization and DevOps. You own the entire technology infrastructure life cycle: Engineer and support software and systems infrastructure. Introduce new foundational technologies that advance our software engineering capabilities to the next level. Collaborate with our software development teams on support issues and improvements to our infrastructure tools, processes, and software. Act as a conduit between our application development teams, and IT, network security, and other stakeholders to align priorities and translate business requirements into technical designs. Improve systems infrastructure reliability. Gather and analyze metrics from operating systems and applications to assist in performance tuning, fault finding and business continuity planning. Design, plan and implement solutions in an entrepreneurial spirit. What You Bring Programming Knowledge – you have an undergraduate, graduate, or post-graduate degree in a computer-related field OR exceptional programming skills gain
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu
About the Team This team builds an AI-native predictive analytics platform that embeds ML and AI-driven insights directly into production Go-To-Market workflows — powering real-time decisions at scale. The team owns the full stack from distributed data pipelines and backend services to ML and AI-powered capabilities that ship directly to customers. We are building toward a model where AI components are first-class runtime dependencies, not bolt-on features. Agentic AI development is a core part of how we increase engineering velocity and deliver customer value. This is a team that ships daily, iterates constantly, and treats speed as a capability to be deliberately improved. The Role This is a high-ownership, builder-first Sr. Software Engineer role. You will design, build, and ship AI-integrated data systems from concept through production — owning outcomes end-to-end, including deployment, monitoring, cost, and business impact. We are seeking a candidate who views AI tooling as a fundamental force multiplier in their daily engineering process. This position is central to our transition into an AI-native function, requiring an individual capable of making decisive, pragmatic architectural choices on reversible matters to maintain momentum. We need an experienced builder of production-grade, data-centric systems who is obsessed with delivering customer value and possesses a deep, curious enthusiasm for the transformative potential of AI. What You Will Build AI-Native Systems Development. Design, build, and own scalable data and ML pipelines, backend services, and AI-powered capabilities that are part of the platform's production decision-making layer. AI and ML components are runtime dependencies in this role — not research projects or experiments. Candidates will have strong back end and data engineering skills to thrive in this space. Daily Shipping. Decompose complex work into safely mergeable increments and ship them daily. Treat large, multi-day pull requ
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The CX Intelligence team is part of Coinbase’s Enterprise Applications and Architecture org and builds the customer-facing and internal CX experiences that connect the Help Center, chatbots (CBCB), and agent workflows. The team owns the multi-agent platform that powers Coinbase chat, Help Center, and agent tooling, partnering closely with Conversation Design, CX, and engineering teams to deliver secure, compliant, and scalable AI-powered support. Our work helps customers get answers faster while enabling agents to resolve cases more effectively. We are hiring an IC4 Machine Learning Engineer to help evolve our conversational ecosystem by building a seamless hybrid vendor-internal chatbot experience. You will contribute to the design and implementation of a unified orchestration layer that coordinates interactions between vendor AI, internal multi-agent systems, and human participants. This role is ideal for someone who enjoys solving complex ML systems problems, building reliable handoff logic across LLM frameworks, and shipping AI-enabled products that are measurable and scalable. What you'll do: Build and improve the orchestration layer that manages state transitions, context sharing, and intent routing across vendor and internal LLM frameworks in a distributed conversational environment. Develop production-grade Python services that bridge advanced AI and ML capab
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Staff Software Engineer (EAA) The Customer Experience (EAA-CX) sub-organization within Coinbase's Enterprise Applications and Architecture (EAA) group builds and operates the secure, Tier-1 platforms that power internal investigative and support workflows for approximately 7,000 specialists across CX, Compliance, Legal, Global Investigations, and Trust and Risk. As a Staff Software Engineer, you'll operate as a technical anchor across 8 globally distributed EAA-CX engineering squads, identifying high-leverage AI and developer-experience opportunities, architecting solutions on Coinbase's production AI infrastructure, and driving platform excellence that raises Tier-1 engineering quality across the portfolio. What you'll do: Own the EAA-CX AI maturity and developer-experience roadmap in partnership with engineering leadership, embedding with squads to surface opportunities, prototyping on Coinbase's existing AI platform, and driving adoption through working demos and feedback loops. Lead platform excellence workstreams that improve reliability, observability, change-failure rate, and operational readiness for Tier-1 systems, partnering with engineering managers to ensure practices persist on the teams. Define build-versus-buy strategy for developer-experience and AI tooling across EAA-CX, evaluating what to build internally, adopt from partner teams, procure from vend
We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training Performance Engineer, you’ll drive efficiency improvements across our distributed training stack. You’ll analyze large-scale training runs, identify utilization gaps, and design optimizations that push the boundaries of throughput and uptime. This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale. You’ll help ensure that our clusters are running at peak performance, enabling OpenAI to train larger, more capable models with the same compute budget. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Profil
OpenAI’s charter calls on us to ensure the benefits of AI are distributed broadly and safely. Our Health AI team focuses on expanding access to high-quality medical expertise and aims to set a high standard for deploying AI responsibly in high-stakes domains. Improving health will be one of the defining impacts of AGI. Today, millions of people lack access to reliable medical information, and clinicians around the world face increasing time and resource constraints. We are building AI systems that support patients, clinicians, and health workers, while meeting the highest standards for safety, reliability, and privacy. We are seeking full stack software engineers to help build and scale products used by consumers and care providers globally. You will work closely with product, design, and research teams to ship real systems in a fast-moving, high-impact environment. In this role, you will: Design and build scalable fullstack systems for consumer and enterprise health. Own end-to-end feature development—from early design and implementation through deployment, monitoring, and iteration. Build and maintain data pipelines and services that meet strict privacy, security, and compliance requirements (e.g., HIPAA). Collaborate closely with researchers and safety teams to integrate reliability, evaluation, and guardrails into production systems. Debug, optimize, and harden systems to support high availability, performance, and global scale. Take ownership of ambiguous problems and drive them to practical, high-quality solutions. You might thrive in this role if you: Are deeply motivated by improving health outcomes and expanding access to medical expertise. Are a strong engineer who enjoys building durable, well-designed systems. Have 5+ years of experience writing maintainable, production-quality code. Can operate with high agency—owning problems end-to-end with minimal supervision. Enjoy working in fast-moving, cross-functional teams with engineers, product managers, desi
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Senior Manufacturing Engineer serves as the senior technical authority for assembly process and equipment design within the Industrialization team. This position designs the assembly lines, fixtures, and material flow systems that convert KARNO prototype and development builds into repeatable, scalable production operations — and then systematically removes waste, labor content, and variation from those operations. Working closely with industrialization leadership, design engineering, production, quality, and supply chain, the Senior Manufacturing Engineer owns the most complex assembly value streams end to end: line and station design, fixture and tooling design, material presentation and handling, process qualification, and continuous waste reduction. This position sets the technical standard for how assembly processes are designed and documented at Hyliion, provides mentorship and design review for other manufacturing engineers, and is accountable for measurable improvement in cycle time, first-pass yield, labor content, and ergonomics across the assembly areas. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Assembly Line and Cell Design : Design assembly lines, cells, and workstations for KARNO core, module, and subassembly operations. Establish work sequencing, balance work content to takt, define station layouts and footprints, and design lines that accommodate planned rate increases rather t
The MongoDB Atlas team is a diverse group of contributors working together to help our users manage MongoDB at global scale. We are responsible for MongoDB Atlas: our database as a service offering and fastest growing product which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. We're seeking a Senior Engineer to join the Atlas Identity and Access Management (IAM) team. IAM is a platform and a product team. We serve internal engineers by providing them a secure and durable suite of services, and we serve external customers by providing them user facing features and products. We are the owners of Atlas’ authentication (OAuth, SSO, Federated Identity) and authorization (RBAC, ABAC) systems, along with many others. The IAM team’s mission is to enable customers to securely build their applications with Atlas through our best in class user experience. We are looking to speak to candidates who are based in New York City, NY for our hybrid working model. Role Responsibilities Design, architect, build, and deliver core pieces of IAM Lead projects from specification to delivery Mentor and grow other team members Improve our codebase, best practices, and design principles Define your top priorities and focuses, communicate them, and execute against them Lead and contribute to complex technical projects and initiatives Candidate Profile 5+ years experience of software engineering, primarily focused on backend systems Proficient in a modern compiled programming language (Java, Go, C#, C++, etc.) Willingness to learn JavaScript and/or TypeScript along with modern frontend technologies (React, Redux, etc.); prior experience a plus Excellent communication skills, both written and verbal Desire to collaborate with colleagues and mentor fellow engineers Is curious, collaborative, empathetic, and intellectually honest Has a passion for problem solving and learning new things in the domains of computer science and software engineering Expe
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. The Infrastructure product organization develops New Relic infrastructure instrumentation agents, next generation data processing and management services, vulnerability management, and security testing capabilities for on-prem and cloud customers. We work with data at a scale using a diverse tech stack (Go, Java, JavaScript, React GraphQL, Kubernetes, many public cloud web services, and more). As a senior backend engineer, you will help us build and extend next generation solutions such as a control plane for customers to manage their data pipelines at scale. New Relic is looking for engineers who are interested in building a brand-new observability experience. This high-impact engineering position is a phenomenal opportunity to own and build a set of next generation services and capabilities for the company. We are searching for a motivated engineer who is ready for a career-defining role in their next opportunity. We look forward to talking with you! What you'll do ● Design, Build, maintain, and scale back-end services and their support tools. ● Participate in architectural definitions with a high degr
Location Details: Melbourne, Victoria, Australia At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... GoDaddy is empowering everyday entrepreneurs around the world by providing the tools, insights, and support they need to succeed online. Our mission is to help customers turn their ideas and personal initiative into success. We provide everything needed to build, grow, and manage their businesses online. GoDaddy's Registry team is looking for a Software Development Engineer to join a collaborative Agile team focused on building products and services that power the Registry system. In the role of a Software Development Engineer, you will contribute to user stories. You will build and operate services at the core of the internet. You will take full responsibility for the solutions and systems you build. This is primarily a backend-focused role with opportunities to contribute across the full stack, including UI development. You'll work with modern cloud technologies, scalable distributed services, and AI-assisted engineering tools while building reliable and highly available solutions for GoDaddy Registry! You'll collaborate with engineers across the team, continuously improve code quality and engineering practices, and gain exposure to technologies across the AWS and Java ecosystem. What you'll get to do... Full-stack development: Participate throughout the entire development stack, from UI components through backend services and persistence layers. Build clean, maintainable, and production-ready code with a strong focus on quality and reliability. Cloud infrastructure and applications: Design, develop, maintai
Get new distributed systems engineer jobs by email
Daily job updates · Unsubscribe anytime