Jobiba hiring network

Senior Software Engineer Distributed Systems Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior software engineer distributed systems jobs. Use filters to narrow by work mode, employment type, experience and date posted.

R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's Cache team is building a next-generation caching solution designed to deliver sub-millisecond average latency, horizontal scalability, and high efficiency—all at a drastically lower cost. Our ultimate vision is to shape a caching infrastructure capable of supporting 1 billion Daily Active Users while reducing costs by 90%. We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to solve Roblox's ever-growing caching challenges. You will report directly to the Engineering Manager for the Cache team. (Check out our recent engineering blog post here to learn more about the team's latest work!) You will: Lead the architectural transition to a next-generation, multitenant caching service built on ValKey, ensuring strict data, resource, and failure isolation for all tenants. Drive systemic optimizations to mitigate head-of-line blocking, manage hot keys, and maximize CPU and memory utilization across physical machine clusters. Design and build robust frameworks to a

redisawskubernetes
View job →
NR
17 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity We are a global team of innovators shaping the future of observability. Our intelligent platform gives customers real-time insight into complex systems so they can innovate faster and operate reliably in an AI-first world. If you’re excited by high-throughput distributed systems and want to contribute to one of the largest and fastest-growing observability platforms, we’d love to hear from you. Join a backend engineering team focused on building and operating JVM-based services that ingest, process, and serve massive volumes of telemetry data. You’ll work on high-scale, low-latency systems that power mission-critical observability features used by engineers worldwide. What you'll do Design, build, and operate JVM-based microservices (primarily Java and Kotlin) with a focus on performance, scalability, and reliability. Own services end-to-end: architecture, implementation, deployment, monitoring, on-call participation, and continuous improvement. Apply strong concurrency and performance practices: asynchronous programming, backpressure, efficient I/O, memory management, and GC tuning. Build and evolve event-driven systems; work with Kafka for streaming, partitioning, consumer groups, and schema evolution.Instrument services for deep observability (metrics, logs, traces), define SLIs/SLOs, and use e Experience with Kafka or similar streaming technologies (topic/partition strategy, consumer lag, idempotency, schema compatibility) strongly preferred. Proficiency w

javarediskubernetes
View job →

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. About the Role The AI Platform Engineering team is looking for a highly motivated and talented engineer who are passionate about continuous learning and excited to grow in a fast-paced, innovative environment. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to a Software Engineering Manager and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. What You'll Do Build the AI Platform Foundation : Lead the design and ownership of the core infrastructure that serves as the backbone for all Smartsheet AI experiences. Focus on building a robust, multi-tenant environment that reduces friction for internal teams, allowing them to deploy reliable and scalable AI features with ease. Standardize the AI Developer Path : Architect high-level abstractions and "Golden Path" APIs that democratize AI development across Smartsheet. By insulating product teams from infrastructure complexity, you will enable them to ship intelligent features with high velocity while guaranteeing safety and consistency at scale. Engineer AI Trust & Safety Systems : Establish the mission-critical monitoring and quality assurance layers that protect Smartsheet customers. By creating rigorous evaluation pipelines, you will ensure every AI-driven feature meets the high bar for safety, data privacy, a

REMOTEpythonvueaws
View job →

We're looking for a Senior Engineer with a strong background in computer science fundamentals, systems design, experience in the Java ecosystem, streaming systems, and data-intensive applications to join our engineering team. In this role, you will be instrumental in designing, building, and optimizing the underlying data structures, algorithms, and database interactions that power our generative AI platform, code generation and migration tools. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation and building a sophisticated data migration suite using a modern technology stack, which includes Java, Spring Boot, Kafka, Debezium, and React.You will work on critical components that ensure the scalability, efficiency, and reliability of our services, collaborating closely with AI researchers, product management and other engineers to design and implement cutting-edge products that solve complex customer challenges.. We are looking to speak to candidates who are based in Sydney for our hybrid working model. The ideal candidate for this role will have 6+ years of engineering experience in backend systems, distributed systems, or core platform development. Proficiency in one or several of Java, Rust, C/C++, and/or Python, with a strong understanding of systems-level programming, memory management, and performance tuning. Extensive experience with streaming data platforms such as Apache Kafka and Change Data Capture (CDC) tools like Debezium Extensive experience with relational data modeling and hands-on experience with at least one SQL database (Postgres, MySQL, etc) Exposure to client-side technologies such as JavaScript and React is a plus Good understanding of algorithms, data structures and their time and space complexity Curiosity, a positive attitude, and a drive to continue learning Excellent verbal and wri

javascriptpythonjava
View job →
N
19 days ago

NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready for their most important workloads. In this role, you help design and run services that sit at the intersection of hardware, security, and large-scale distributed systems. We partner closely with security, silicon, platform, and cloud teams to bring new ideas into reliable production services that people rely on every day. We care about building systems that last, supporting each other, and creating space for learning and experimentation. If you enjoy solving complex problems, keeping services running smoothly, and collaborating with teammates from many disciplines, we would love to talk with you! What you’ll be doing: Your main focus will be on building and managing our core attestation cloud services. Day-to-day responsibilities include crafting APIs and integrations, boosting reliability, and working alongside NVIDIA teams to convert hardware trust mechanisms and standards into production-ready solutions. You will contribute significantly to shaping how customers verify that NVIDIA platforms are secure and prepared for their workloads. Crafting and evolving attestation cloud services, APIs, and SDK/CLI integration points that confirm the integrity of NVIDIA platforms across data center, AI, networking, and partner environments. Improving reliability and operational maturity through SLOs/SLIs, alerting, runbooks, incident response, and safe rollout practices. Crafting resilient service behavior that handles dependency failures, caching challenges, regional issues, customer-side resilience needs, and graceful degradation. Architecting trust-material distribution for certificate status, re

javaawsazure
View job →

Position Overview: We are looking for a Senior Software Engineer to drive technical excellence, architect complex systems, and elevate our engineering team. You will own critical technical decisions, lead major initiatives from conception to delivery, and set the standard for engineering quality across our products. As a senior engineer, you will architect and lead the development of sophisticated AI-enabled features and infrastructure. This includes designing MCP server architectures, building advanced RAG systems, implementing agentic AI workflows, and establishing patterns that scale across our product portfolio. You will combine deep technical expertise in both traditional software engineering and AI/ML to deliver production-grade solutions. What You'll Do Lead the design and implementation of AI integration infrastructure (MCP servers, orchestration layers, API gateways) Build sophisticated AI features including advanced RAG systems, agentic workflows, and multi-step reasoning Establish AI engineering best practices, security patterns, and quality standards Lead technical initiatives from requirements through production deployment Make critical architectural decisions balancing performance, scalability, cost, and maintainability Design AI evaluation frameworks and implement quality benchmarks Debug and resolve complex production issues across traditional and AI systems Required Qualifications Tech Stack Core: Node.js, React, TypeScript, AWS, PostgreSQL, MSSQL, Docker AI & Integration: Python, MCP, AWS Bedrock, LangGraph/Semantic Kernel, Vector Databases, RAG Core Technical Skills 5+ years professional development with proven track record of delivering complex systems Strong Node.js and JavaScript/TypeScript expertise Advanced React and frontend architecture skills Extensive AWS architecture experience Expert PostgreSQL database design, optimization, and performance tuning Deep understanding of microservices, distributed systems, and

javascripttypescriptpython
View job →
H
22 days ago

Roles and Responsibilities Write maintainable/scalable/e?cient code. Work in a cross-functional team, collaborating with peers during the entire SDLC. Follow coding standards, unit-testing, code reviews etc. Follow release cycles and commitment to deadlines. Qualifications & Experience Experience level of 3 to 5 years of experience in very large scale applications. Fair understanding in problem solving skills, data structures and algorithms. Experience with distributed systems handling large amounts of data. Fair understanding in coding skills in Java/J2EE, Web technologies, RDBMS/messaging.

javaaigo
View job →
H
22 days ago

Roles and Responsibilities Write maintainable/scalable/code Work in a cross-functional team, collaborating with peers during the entire SDLC. Follow coding standards, unit-testing, code reviews etc. Follow release cycles and commitment to deadlines. Qualifications & Experience Experience level of 3 to 6 years of experience in very large scale applications. Fair understanding in problem solving skills, data structures and algorithms. Experience with distributed systems handling large amounts of data. Fair understanding in coding skills in Java/J2EE, Web technologies, RDBMS/messaging.

javaaigo
View job →
H
22 days ago

Roles and Responsibilities Write maintainable/scalable/e?cient code. Work in a cross-functional team, collaborating with peers during the entire SDLC. Follow coding standards, unit-testing, code reviews etc. Follow release cycles and commitment to deadlines. Qualifications & Experience Experience level of 3 to 5 years of experience in very large scale applications. Fair understanding in problem solving skills, data structures and algorithms. Experience with distributed systems handling large amounts of data. Fair understanding in coding skills in Java/J2EE, Web technologies, RDBMS/messaging.

javaaigo
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $170K – $240K/yr
22 days ago

Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible. What You Will Be Doing Build observability tools and platforms, including: metrics, logging, distributed tracing, dashboarding, alerting, application performance management Build with modern tools and languages like Go, Open Telemetry and Kubernetes Participate in on-call rotation and ensure uptime of services Create runtime tools/processes that optimize cloud triaging and limit downtime Define best practices around making our systems and services measurable Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies. We expect successful candidates to be coding a majority of their time Qualifications We Need Strong Computer Science fundamentals 5+ years industry experience building and maintaining high-quality software, especially software other engineers use You apply a product mindset to infrastructure systems and feel accomplished enabling others Desire to be a great teammate and have fun at work Strong sense of craftsmanship, and a healthy academic curiosity Qualifications We Want (also, skills you’ll learn!) Experience building systems for data analytics Distributed systems monitoring and profiling skills Knowledge of cloud application security models Administered cloud service infrastructure (GCP, AWS, Azure) Startup experience Additional Job details Additional Job details The base salary range for this position is $170k - $240k annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience. Base pay is one part of the Total Package that is provided to compensate and recognize e

pythonsqlaws
View job →
DC
22 days ago

Here's a summary of the role: Do you love building scalable cloud platforms and solving complex engineering problems with modern technologies? As a Senior Software Engineer at Diligent, you'll design and deliver high-performing , serverless applications that power our global SaaS platform. You'll work extensively with TypeScript, Node.js, AWS, and event-driven microservices, owning services from design to deployment and production monitoring. This is an opportunity to influence technical decisions, mentor engineers, and explore how AI can transform software development and engineering productivity. If you're passionate about cloud-native architectures, distributed systems, and building software that scales to millions of users, we'd love to meet you. Here's a breakdown of what you'll do (not all of it, just the important stuff): Design and build scalable backend services and event-driven microservices using TypeScript and AWS. Develop secure APIs and integrations that power reporting, analytics, and dashboard experiences. Build and maintain serverless solutions using AWS services such as Lambda, EventBridge , SQS, and DynamoDB. Drive engineering excellence through testing, observability, automation, and production readiness practices. Contribute to infrastructure-as-code and CI/CD pipelines using AWS CDK and modern DevOps practices. Mentor engineers, participate in architecture discussions, and champion the use of AI tools to improve development efficiency. These are the essentials you'll need to get an interview: 6-8 years of professional software engineering experience. Strong experience with TypeScript, Node.js, and modern backend development patterns. Hands-on experience building cloud-native applications on AWS. Strong understanding of serverless architectures and event-driven microserv

typescriptreactnode.js
View job →
G
22 days ago

About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team An exciting opportunity to join a new team within the Software Operations group. The Build Engineering team is a new function within Software Infrastructure, which focuses on the overall process of building and integration of the Machine Learn ing S oftware S tack. You will work closely with the QA and development teams to get an understanding of how our ML SW stack is built, helping to ensure good build practices, and proving that the stack works together and is reproducible in secure, sandboxed environments. Responsibilities and Duties Developing our internal t

pythondockerlinux
View job →
DU
DoorDash USA
📍 San Francisco• Full-time• From $1.6M/yr
22 days ago

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte

pythonjavasql
View job →
L
Lyft
📍 Toronto• Full-time• From C$136K/yr
29 days ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and

pythonawsdocker
View job →
C
Coder
📍 United Kingdom• Full-time• Remote
1mo ago

As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building integrations across multiple model providers. Experience with AW

REMOTEtypescriptreactaws
View job →
🔔

Get new senior software engineer distributed systems jobs by email

Daily job updates · Unsubscribe anytime