Jobiba hiring network

Distributed Systems Engineer Jobs

1,306 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production

awsrestai
View job →
A
1mo ago

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's need to detect and respond to security events across its production and corporate environments is growing as the company scales. We're looking for a Senior Detection and Response Engineer to own detection engineering and to lead incident response when it counts, coordinating the response and driving it to resolution. This is a high-ownership role with real room to shape how detection and response works at Anyscale. You will own the detection pipeline, the response runbooks, and incident response, reporting to the Head of Security and partnering with engineering. This role is based in India. In your first year, success looks like strong detection coverage across our cloud, endpoint, and runtime telemetry, a working correlation and alerting pipeline, and incident response runbooks that have been exercised in practice. What You'll Do Own and build detection coverage across cloud, endpoint, and runtime telemetry. Own a centralized correlation and alerting capability that turns telemetry into actionable detections. Own incident response: runbooks, escalation paths, and coordination during an incident, across corporate and production environments. Drive detection of anomalous activity across the environments

awsazurekubernetes
View job →
C
Coder
📍 United Kingdom• Full-time• Remote
1mo ago

As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building integrations across multiple model providers. Experience with AW

REMOTEtypescriptreactaws
View job →
C
1mo ago

As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.

REMOTEtypescriptreactaws
View job →

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team focused on transforming how Mastercard's payment systems are built, scaled, and operated. As a Senior Software Engineer, you will lead the design and development of cloud-ready applications, microservices, and distributed systems that support large-scale payment processing platforms while helping advance modernization, automation, and engineering excellence across the organization. In this role, you will contribute to software architecture decisions, drive technical design discussions, and partner with engineers to deliver scalable, resilient, and maintainable software solutions. You'll have the opportunity to solve complex technical challenges, mentor other engineers, and influence how software is designed, developed, tested, and supported across critical technology platforms. What You Will Do •Design software solutions and contribute to software architecture decisions that support scalability, maintainability, and operational excellence. •Translate complex product requirements into technical designs and implementation plans. •Lead development of modular, extensible, high-performance applications. •Design and implement comprehensive unit, functional, and integration testing strategies. •Analyze, optimize, and improve application performance, scal

javaairecruitment
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity The Streaming Services team is responsible for routing and replicating New Relic's platform data through globally distributed cloud environments in a high-throughput, low-latency, and cost-aware streaming pipeline. Our streaming services process and route over a billion messages per minute across different continents, regions, and cloud providers, serving New Relic’s customer experience by making it easy for development teams to prioritize reliability. You will collaborate with this globally distributed team to develop expertise and best practices for New Relic development teams to operate highly reliable streaming services. What you'll do Own, build, maintain, and scale our streaming services and their support tools. Participate in an on-call rotation and bake stability into everything, continually seeking automation opportunities for built-in reliability. Participate in architectural definitions with a high degree of innovation and creativity. Own and improve your team processes. Develop automation and tooling to make our services more scalable and reliable. Use available innovation time to bring your creative ideas to life. This role requires Experience developing back-end services that use Flink, Kafka, or other streaming platforms. Large scale is a plus. Experience in writing software in Java, and you are not afraid of adapting, learning, and working with different languages and frameworks. Experience with distributed systems, concurrency, and scaling in

javakubernetesai
View job →
A
1mo ago

Job Requisition ID # 26WD97363 Position Overview We are seeking a Principal Software Engineer – Backend to join Autodesk’s Enterprise Data Management (EDM) organization within the COO-GET Engineering group. This is a senior individual contributor role operating at the Principal (P4) level , expected to drive technology direction for large, complex, and business-critical backend and distributed systems . This role is anchored in backend software engineering excellence : designing, building, and evolving scalable services, APIs, and event-driven systems that operate at enterprise scale. As a Principal Engineer, you will work with high autonomy and ambiguity , shape long-term architecture, and influence multiple teams and domains. Familiarity with data engineering concepts is valuable, but backend systems, service design, and distributed systems are the core competencies. You will function as a technical authority and force multiplier—guiding design decisions, setting standards, and ensuring Autodesk’s core data services are reliable, resilient, and evolvable over time. Responsibilities Provide principal-level technical

REMOTEpythonawsai
View job →
N
Nvidia
📍 Bengaluru, India
1mo ago

We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations. What You Will Be Doing: Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale. Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams. Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals. Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation. Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams. Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness. Provide technical leadership and mentorship, influence engineerin

pythonkubernetesai
View job →
M
Mongodb
📍 Toronto• Full-time• From C$158K/yr
1mo ago

MongoDB Search and Vector Search allows users to execute complex search queries and build RAG applications using the MongoDB Query Language. Our users are free to focus on relevance and data retrieval instead of the machinery needed to search data at scale. We are looking to speak to candidates who are based in Toronto for our hybrid working model. Candidate Profile: 5+ years of hands-on experience designing, building, testing, and maintaining industrial-strength backend software in a complex codebase Proficient in modern programming languages and techniques Experienced in developing distributed systems, cloud services, and SaaS products Excellent verbal and written technical communication skills; enthusiasm for collaborating closely with colleagues and mentoring other engineers A growth mindset and the desire to learn quickly through taking on challenges, reflecting on outcomes, and incorporating feedback A strong sense of ownership over their work, from initial design all the way through maintaining code in production You will: Build and design our integrated search platform, written in Java Work with a collaborative team that prioritizes sound technical decision-making and building systems that our customers love and that we are proud of as engineers Lead projects and own subsystems Help determine the team’s roadmap and the architecture of our system Success measures: In 3 months you’ll have contributed to the development of an existing project and completed several improvements or bug fixes In 6 months you’ll be reviewing code and project designs, and be an active participant in team meetings In 12 months you’ll have a thorough understanding of the systems the team owns and have led a project. You’ll have had a positive impact on our code, product, and team processes About MongoDB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to cr

javamongodbaws
View job →
O
1mo ago

About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond to operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will design and develop systems that support the model lifecycle in production, including deployment orchestration, configuration management, operational automation, reliability, and capacity management. You will help transform complex operational processes into scalable platform capabilities that enable teams across OpenAI to deploy and manage models with greater confidence and less manual effort. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build and evolve the platform used to deploy, configure, and manage models across ChatGPT. Develop systems for deployment orchestration, model rollouts, operational visibility, and production readiness. Create abstractions and tooling that simplify complex infrastructure and improve the developer experience. Automate operational workflows, including incident detection, diagnosis, mitigation, and recovery. Improve the reliability, scalability, and efficiency of model deployments and the infrastructure that supports them. Build systems that support capacity planning, resource allocation, and infrastructure utilization. Partner with research, infrastructure, and product engineering teams to identify common chal

REMOTEpythonawsrest
View job →
C
Coder
📍 United Kingdom• Full-time
1mo ago

As an Engineering Manager on Coder’s Agentic Engineering team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction, grow the team, and keep execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Agentic Engineering organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development enviro

typescriptreactaws
View job →
V
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta’s new Enterprise Resilience team is being formed to support the next era of growth by powering services that are reliable, scalable, and resilient by design. As our customer base expands and our systems scale, we need a dedicated group focused on partnering closely with product engineering teams to build and operate robust distributed systems across all of Vanta’s environments, including our new FedRAMP deployment. In this role, you’ll help define the foundations of reliability at Vanta including shaping best practices, building core infrastructure, and guiding teams as they design services that perform consistently for customers. This team will have a broad and deep impact across product engineering. Your work will influence how every Vanta engineer builds, deploys, monitors, and maintains their services, whether for our commercial environment or regulated customers with more stringent requirements. You’ll develop tools and frameworks that make it easier to detect and remediate issues, improve operational readiness, and support feature development that meets the needs of increasingly large and complex enterprise customers. Vanta engineers design and develop new product functionality and infrastructure using modern frameworks and tooling, including TypeScript, React, Node.js, MongoDB, GitHub Actions, and AWS services such as Fargate and ECS. If you're excited to help define a new function, raise the reliability bar across an entire engineering organization, and build systems that scale with Vanta’s growth, we’d love to meet you. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll

typescriptreactnode.js
View job →
C
Clickup
📍 United States• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Come join the Hierarchy Squad to build and optimize a set of tier 1 services that are at the core of ClickUp’s all-in-one productivity platform! Hierarchy owns the ClickUp filesystem that is at foundation of everything we build. We are searching for passionate engineers with exceptional critical thinking skills and technical prowess to help maintain, develop and scale high throughput services to support the company through accelerated growth. The ideal candidate will find the idea of architecting distributed systems that are reliable and highly performant exciting, has a knack for debugging and writing complex code and is excited by the challenge of building and maintaining a platform that every team in the company integrates with. Our key technologies include Typescript, Postgres, Kafka, NestJS running on Amazon Web Services. If this sounds interesting to you, we'd love to have you join our team! Responsibilities: - Develop and maintain robust, scalable backend systems using Node.js (Express and NestJS). - Collaborate with engineers, designers, and product managers to drive projects forward. - Tune and optimize database queries for maximum efficiency and performance. - Optimize and improve existing code for better performance and user experience. - Troubleshoot and debug issues, ensuring smooth operations. - Share your knowledge and expertise to foster a culture of learning and growth. Requirements: - 8+ years of professional experience building backend services for SaaS products. - Proven track record of building and scaling backend systems. - Expertise in relational database query optimizations (pre

typescriptreactnode.js
View job →
S
Sentry
📍 Toronto• Full-time• C$162K – C$420K/yr
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed

aigorust
View job →
S
Sentry
📍 San Francisco• Full-time• $155K – $400K/yr
1mo ago

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed

aigorust
View job →
🔔

Get new distributed systems engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More distributed systems engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.