We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. Database Observability is a critical pillar of New Relic's platform strategy. We are looking for an experienced Engineering Manager (M3) to lead a senior, high-performing team building next-generation Database Observability products. Your team will own the entire lifecycle of critical telemetry data flows from lightweight database agents and high-throughput ingestion pipelines to intelligent DB recommendation engines and autonomous DB AI Agents. You will lead a team that includes senior and Lead-level engineers with deep domain expertise in distributed systems and AI. Your primary value will come from setting strategic technical direction, enabling their best work, and fostering a high-accountability culture while partnering closely with Product and Design to deliver features that directly drive New Relic's Database Observability. What you'll do Manage a full-stack engineering team (6–8 engineers) spanning backend systems, database telemetry, agent engineering, and UI workflows. Own end-to-end delivery sprint planning, roadmap execution, system quality, and operational excellence for critical database ingesti
Jobiba hiring network
Software Engineer Distributed Systems Manager Manager Manager Jobs
15 active opportunities · Updated for September 2026
Fresh results
15 shown
Explore current software engineer distributed systems manager manager manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. The Realtime Infrastructure team is responsible for building and maintaining some of Discord’s highest scale and most critical services. Those systems are at the core of our text chat infrastructure and facilitate the dispatching of every update to our users sessions. This role will have a significant impact on Discord’s overall reliability and performance. It will also help our product teams build new features on top of our infrastructure. This team is small but critical, and its work has a direct impact on Discord's success and ability to scale. This role reports to the Senior Engineering Manager of Realtime Infrastructure. What You'll Be Doing Build and operate large-scale, reliable and performant distributed systems. Collaborate with product teams to create new features. Ensure Discord “just works”. Write code but also manage our infrastructure. Work with a talented team of engineers who have built one of the largest communication platforms in the world. What you should have 2+ years of experience writing and designing backend systems. Experience solving complex distributed system problems. Experience operating and maintaining critical tier 0 services. Knowledge of monitoring and alerting best practices. Familiar with open source software, and not afraid to dig into the source code of a library to find the answer you’re looking for. Bonus Points Experience with Elixir or Rust. Experience working with systems deployed in a cloud environment (GCP, AWS, etc.) Knowledge of devops tools like Salt,Terraform or k8s. You have built or contributed to open source projects. You are a Discord power user and hav
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Ray aims to provide a universal API for building distributed applications (e.g. a machine learning pipeline of feature engineering, model training, and evaluation). Data is usually a core element connecting these different stages, and therefore plays a critical role in Ray’s usability, performance, and stability. We are looking for strong engineers to build, optimize, and scale Ray’s Datasets library and data processing capabilities in general. About the Ray Data team: The Ray Data team currently develops and maintains the Ray Datasets library, which is already powering critical production use cases (e.g. large scale data compaction at Amazon , and ML pipeline at Alibaba ). Ray Datasets is a Python library built on top of Apache Arrow and Ray Core (Ray’s C++ backend), and the Ray Data team interacts closely with Ray Core components including the scheduler and the memory & I/O subsystems. The Ray Data team also works closely with Ray’s ML libraries including Train, RLlib, and Serve. A snapshot of projects you will work on: - Performance of Ray Datasets at large scale (leveraging Arrow primitives, optimizing Ray object manager, etc.) - Integration with ML training and data sources - Stability an
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's Cache team is building a next-generation caching solution designed to deliver sub-millisecond average latency, horizontal scalability, and high efficiency—all at a drastically lower cost. Our ultimate vision is to shape a caching infrastructure capable of supporting 1 billion Daily Active Users while reducing costs by 90%. We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to solve Roblox's ever-growing caching challenges. You will report directly to the Engineering Manager for the Cache team. (Check out our recent engineering blog post here to learn more about the team's latest work!) You will: Lead the architectural transition to a next-generation, multitenant caching service built on ValKey, ensuring strict data, resource, and failure isolation for all tenants. Drive systemic optimizations to mitigate head-of-line blocking, manage hot keys, and maximize CPU and memory utilization across physical machine clusters. Design and build robust frameworks to a
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. About the Role The AI Platform Engineering team is looking for a highly motivated and talented engineer who are passionate about continuous learning and excited to grow in a fast-paced, innovative environment. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to a Software Engineering Manager and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. What You'll Do Build the AI Platform Foundation : Lead the design and ownership of the core infrastructure that serves as the backbone for all Smartsheet AI experiences. Focus on building a robust, multi-tenant environment that reduces friction for internal teams, allowing them to deploy reliable and scalable AI features with ease. Standardize the AI Developer Path : Architect high-level abstractions and "Golden Path" APIs that democratize AI development across Smartsheet. By insulating product teams from infrastructure complexity, you will enable them to ship intelligent features with high velocity while guaranteeing safety and consistency at scale. Engineer AI Trust & Safety Systems : Establish the mission-critical monitoring and quality assurance layers that protect Smartsheet customers. By creating rigorous evaluation pipelines, you will ensure every AI-driven feature meets the high bar for safety, data privacy, a
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With over half a billion rides and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Marketplace, Mapping, Fraud, Growth and beyond. Building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives requires robust, scalable software systems operating at massive scale. Our highly motivated Software Engineers work on these challenging problems and build the systems that directly impact various aspects of our core business. If you are a critical thinker with experience building software systems, passionate about solving business problems through well-crafted code and working in a dynamic, creative, and collaborative environment, we are searching for you. As an associate software engineer, you will be designing, building, and launching the services that power the platform's core products. Compared to similarly-sized technology companies, the set of problems that we tackle is incredibly diverse. They cut across transportation, distributed systems, backend services, mapping, personalization, and real-time infrastructure. We are hiring motivated engineers across each of these areas. We're looking for someone who is passionate about solving problems with code, building reliable and maintainable systems, and is excited about working in a fast-paced, innovative, and collegial environment. You will report to a Software Engineering Manager. Responsibilities: Partner with Engineers, Data Scientists, Product Managers, and Business Partners to build software for business and user impact Perform technical analysis and build proof-of-concept prototypes to explore and propose solutions to both new and existing problems Design and develop software components, services, and APIs Write production quality code to launc
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. WHAT WE OFFER Paid, full-time internships Post-internship career opportunities (full-time or part-time) Exposure to a fast-paced, fun, and inclusive culture A chance to work with world-class security experts on challenging, high-impact projects Opportunity to provide meaningful contributions to a real system used by customers High level of access to supervisors (manager and mentor), detailed direction without micromanagement, feedback throughout your internship, and a final evaluation Treated as a full member of the Snowflake team: company meetings and activities, flexible hours, casual dress code, swag, and much more Warsaw office perks: catered lunches, access to recreational games, happy hours, company outings, and more WHAT WE EXPECT Must be actively enrolled in an accredited college/university program during the internship period Desired majors: Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field Required coursework: algorithms, data structures, and operating systems Recommended coursework: information security, cryptography, cloud computing, database systems, or distributed systems Duration: 4 months + recommended Strong programming skills in Python or Java Solid knowledge of data structures and algorithms; familiarity with
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Hybrid-or-Remote: This position may be a hybrid or fully remote position, as decided by your manager. If designated as hybrid, you’ll divide your time between working remotely from your home and an office location, so you should live within commuting distance. If designated as remote, you’ll be working remotely from your home and may occasionally visit a GoDaddy office to meet with your team for events or meetings. Your hiring manager can share more about this role’s hybrid or remote designation. Join our Team We're looking for a Software Engineer II to join the team responsible for building and supporting services that connect GoDaddy's eCommerce systems with our Data Platform. This role is ideal for an engineer who enjoys backend development, wants to deepen their cloud and distributed systems knowledge, and is excited about learning from experienced teammates while building software that powers millions of customer interactions. You'll work as part of a collaborative Agile team, contributing to the design, development, testing, and operation of modern Java-based services running in AWS. What you'll get to do... Develop and maintain backend services using Java and Spring Boot Contribute to projects that help move and process data between GoDaddy systems. Write clean, well-tested, maintainable code Participate in code reviews and engineering discussions Troubleshoot and resolve issues in your team's services Collaborate with engineers, product managers, and other stakeholders to deliver customer value Learn and apply cloud-native development practices in AWS Contribute ideas for improving applications, processes, and team practices Your experience should include... 1+ years of professional software development
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to identify, understand, and tackle issues, analyze performance, and maximize their software and infrastructure. The Developer Platform organization is looking for a Software Engineering Manager to join our group in India. This organization develops and maintains the internal developer platform, security, compliance and audit reporting and all processes and controls that enable New Relic’s fleet of engineers to effectively release our products to customers. Developer Platform teams work across all engineering teams to implement critical projects to support our business. Our Engineering Managers are a select group of technology leaders and mentors who shape the solutions we bring to the market and foster an inclusive environment that brings out the best in people. If you have a strong technical background but you are also a person with great mentoring skills and that knows how to collaborate with multiple teams to ship great software, we would like to hear from you. What you’ll do Treat the internal developer platform as a product. Define roadmaps, manage backlogs, and prioritize features based on internal engineer needs and business objectives. Drive and contribute to our culture of operational and business efficiency:
About the Team The Storage organization builds and operates the online stateful systems and abstractions that DoorDash Engineering depends on: reliable, efficient, secure, and easy to use. Within Storage, the Distributed Caching team owns every caching offering at DoorDash end to end, including ElastiCache (Redis/Valkey), Boulder (our KVRocks-based key-value store for high-QPS feature serving), Entity Cache (read Bill Shen’s engineering blog post, “ High-Performance Proxy Cache for DoorDash Services ”), and the Distributed Lock Service, plus the smart clients (asgard-redis, valkey-go) that sit in front of them. These systems back critical product surfaces across DoorDash, Wolt, and Deliveroo: the team runs roughly 400 ElastiCache clusters serving hundreds of millions of GET requests per second in aggregate, and Boulder, our offline-to-online feature store, serves billions of feature lookups per second at peak. About the Role The team owns provisioning of clusters and the smart clients that sit in front of them, baking in sensible defaults so that other engineering teams get a turnkey caching solution instead of having to run their own. You'll help drive Boulder's evolution to scale further, improve cost efficiency, enhance performance, and support real-time updates; re-platform the Distributed Lock Service onto a strongly consistent backend; and build the self-serve tooling and recommendation engine that let customers describe a workload (QPS, TTL, payload size, latency profile) and get the right backend without talking to a human. You'll go deep on cache invalidation, replication, sharding, compaction, and failover, while shipping the guardrails, automation, and observability that keep this scale operable by a small team. You must be located in San Francisco, Seattle, or the New York Metro Area for this hybrid position. You will report to the Engineering Manager on the Distributed Caching team within the Storage organization. You’re excited about this opportunity b
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model. This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin. Responsibilities Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems Participate in a 24/7 on-call rotation to resolve critical issues Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice You may be a good fit if you Have 6+ years of experience in software development and operating distributed systems Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests) Have deep experience using and extending containerization technologies, preferably Kubernetes Have a solid understanding
MongoDB Technical Services Engineers use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and to solve their complex MongoDB problems. Technical Service Engineers are experts in the entire MongoDB ecosystem - database server, drivers, cloud and infrastructure. This also includes services such as Atlas (database as a service), or Cloud Manager (which helps customers with automation, backup and monitoring of their MongoDB systems). Our engineers combine their MongoDB expertise with passion, initiative, teamwork and a great sense of humor to help our customers to be successful with MongoDB. We are looking to speak to candidates who are based in Austin for our hybrid working model. Cool things you’ll do You'll be working alongside our largest customers, solving their complex challenges - resolving questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices in running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs - interfacing with our product management and development teams on their behalf. And you'll contribute to internal projects, including software development of support tools for performance, benchmarking, and diagnostics. What you need We consider all candidates with an eye for those who are self-taught, insatiably curious, and multi-faceted. The ideal candidates should have strong technical experience in one (or more) of the following areas Systems administration Distributed systems Network Administration Database architecture and administration Application Architecture Data architecture and design Performance tuning and benchmarking Extra bonus points if you have experience in one or more of Java, Python, Ruby, C, C++, C#, Javascript, node.js, Go, PHP, or Perl If you have an operations background, we prefer experience administering large-scale production environments, i
The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Training & Serving team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: distributed training of foundation models, serving at scale, designing the user experience. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Training & Serving team, directly managing 10+ engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage, infrastructure and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: Previous experience (1+ years) leading software engineering teams, as a tech lead or people manager Strong technician with a mix of backend, data engineer and infrastructure experience who is interested in remaining a hands-on leader Excellent leader with strong
Get new software engineer distributed systems manager manager manager jobs by email
Daily job updates · Unsubscribe anytime