About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in San Francisco to join our global Support Engineering team. This role is designed to provide APAC coverage to our users; with the shift being Sunday through Thursday 4PM-12AM PST. We are architecting the Technical Support engine . We’re looking for an experienced engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, s
Jobiba hiring network
Distributed Systems Engineer Jobs
1,306 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. About the Role The AI Platform Engineering team is looking for a highly motivated and talented engineer who are passionate about continuous learning and excited to grow in a fast-paced, innovative environment. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to a Software Engineering Manager and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. What You'll Do Build the AI Platform Foundation : Lead the design and ownership of the core infrastructure that serves as the backbone for all Smartsheet AI experiences. Focus on building a robust, multi-tenant environment that reduces friction for internal teams, allowing them to deploy reliable and scalable AI features with ease. Standardize the AI Developer Path : Architect high-level abstractions and "Golden Path" APIs that democratize AI development across Smartsheet. By insulating product teams from infrastructure complexity, you will enable them to ship intelligent features with high velocity while guaranteeing safety and consistency at scale. Engineer AI Trust & Safety Systems : Establish the mission-critical monitoring and quality assurance layers that protect Smartsheet customers. By creating rigorous evaluation pipelines, you will ensure every AI-driven feature meets the high bar for safety, data privacy, a
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Staff engineer (L4), Twilio’s Segment team. About the job As a Staff Engineer on the Twilio Segment Data platform/ pipelines team, you’ll build and scale systems that process several hundred thousands of data points per second. You will lead the development of high-scale ingestion and data processing systems You'll be designing, operating and maintaining complex distributed systems, ensuring reliability, performance, and cost-efficiency while querying petabytes of data for our customer data platform (CDP). Responsibilities In this role, you’ll: Design and deliver robust, high-scale routing experiences for the Data platform/ pipelines team for Twilio Segment. Ship features that opt for high availability and throughput with eventual consistency Collaborate with engineering and product leads, as well as teams across Twilio Segment Support the reliability and security of the platform Build and optimize globally available and high
Join the MongoDB Server Query Optimization team, and help us build a world-class distributed open-source query optimizer. Our team plays a crucial role in the experience and performance of data processing. We are responsible for the MongoDB Query Language and the lifecycle of each query, through parsing, optimization and plan selection. We have a presence across the US and Europe including New York, Dublin, Seattle, Palo Alto, and Chicago. We support office-based and remote work and align projects with convenient work hours for each time zone. We have tons of interesting problems to solve with a direct impact on users for transactional, time-series, and analytical workloads. The team is endeavoring to systematically rewrite every major component of our optimization and execution systems. We need your help to design and build the heart of a distributed, flexible schema, document database. This role can be based out of our US offices or remotely in the North America region. Candidate Profile 10+ years of experience in data management systems, distributed systems, or large-scale backend engineering Experience with building production-level code with a large user base, robust design structure and rigorous code quality Degree in Computer Science or similar field, or equivalent practical experience, with strong competencies in data structures, algorithms, and software design/architecture Experience with large code bases written in C++ or another systems programming language. You'll need to trace down defects, estimate work complexity, and design evolution and integration strategies as we rewrite different components of the system A strong foundation in core database internals is essential. While direct experience in query optimization is a massive bonus, it is not a prerequisite. We are also excited to meet candidates with strong backgrounds in compilers, language transpilers, or distributed storage systems Position Expectations Innovate in the area of flexible schema d
The Storage Layer Services team is currently re-architecting the MongoDB Cloud Storage Layer. This is a relatively new team in MongoDB that sits at the heart of the next generation MongoDB Cloud Storage Architecture, and the team is working to build performant multi-tenant distributed storage services both to enhance our existing MongoDB cloud storage architecture and to power more of our customers' use cases more efficiently. Engineering at MongoDB is globally distributed, with a mix of folks being fully remote, hybrid, or in-office. We have a small but growing team that calls Sydney home, and we are looking for a Staff Engineer to join the team working closely with other teams in Sydney and North America. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies great engineering fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. We are looking to speak to candidates who are based in Sydney for our hybrid working model. You’re an ideal candidate if: You have 10+ years of experience in programming, debugging, and performance tuning highly concurrent and/or distributed systems. Especially if you have worked in a systems language (C, C++, Rust, etc) for a number of those years You have a track record as an effective technical leader. You love helping teams be successful at solving vaguely defined problems in iterative and measurable ways. You put the customer first, and don’t hesitate to cross team boundaries in search of the right solution You have a solid grasp of related systems fundamentals, such as cache management, log-based recovery, transactions or performance profiling You’re comfortable reasoning about highly concurrent, asynchronous services — backpressure, tail latency, and the failure modes of replicated state machines You’ve worked on large, highly availabl
We are seeking a Staff engineer to design, build, and operate the internal and external Observability stack for the MongoDB platform. Tens of thousands of customers depend on our Observability stack to monitor their database clusters and to generate actionable alerts to safeguard critical workloads. This is an opportunity to join a team that is responsible for all Observability systems that support metrics, metric visualization, logs, traces, and alerts for MongoDB. We are looking for engineers with high standards, and experience in setting direction and technical leadership for large engineering teams in designing and operating complex distributed systems, with strict SLO on security, durability, availability and performance. As MongoDB Atlas and its supporting infrastructure continue to experience rapid growth, the demand for high-cardinality observability data for internal and external use cases means we need to continually innovate and scale our systems to the next level. For example, MongoDB Observability systems need to handle 10’s of billions of metrics time series, all whilst processing petabytes of logs, traces, and events. Our stack includes VictoriaMetrics, Splunk, Flink, WarpStream/Kafka, Java, Golang Fluentbit. In addition to owning critical components of our observability infrastructure, as a Staff engineer on the team, you’ll also work closely with other SWE, Product and SRE teams to promote and implement best practices in instrumenting and monitoring their services. This is a highly collaborative role, and you will get to own some of the most relied upon internal infrastructure at Mongo. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies low-level systems expertise to build the foundational infrastructure of a popular database, join us! Let’s build a faster, more reliable, and exceptionally observable database system together. W
As a Research Engineer on our team, you will partner with Research Scientists to turn research ideas into working systems, building the data, tooling, and infrastructure that enable rapid iteration, trustworthy evaluation, and a smooth path from prototype to production. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Build and operate multimodal data pipelines, training and evaluation infrastructure, benchmarks, and internal tooling Implement models, run experiments at scale, and profile for reliability, performance, and cost Build simulation environments and replay infrastructure for agent training and evaluation Orchestrate distributed training and distributed RL with Ray, including scheduling, scaling, and failure recovery Establish rigorous automated benchmarks and regression tests for world model predictions, agent performance, and simulation fidelity Collaborate with Research Scientists, Product, and Engineeri
We're looking for a Senior Engineer with a strong background in computer science fundamentals, systems design, experience in the Java ecosystem, streaming systems, and data-intensive applications to join our engineering team. In this role, you will be instrumental in designing, building, and optimizing the underlying data structures, algorithms, and database interactions that power our generative AI platform, code generation and migration tools. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation and building a sophisticated data migration suite using a modern technology stack, which includes Java, Spring Boot, Kafka, Debezium, and React.You will work on critical components that ensure the scalability, efficiency, and reliability of our services, collaborating closely with AI researchers, product management and other engineers to design and implement cutting-edge products that solve complex customer challenges.. We are looking to speak to candidates who are based in Sydney for our hybrid working model. The ideal candidate for this role will have 6+ years of engineering experience in backend systems, distributed systems, or core platform development. Proficiency in one or several of Java, Rust, C/C++, and/or Python, with a strong understanding of systems-level programming, memory management, and performance tuning. Extensive experience with streaming data platforms such as Apache Kafka and Change Data Capture (CDC) tools like Debezium Extensive experience with relational data modeling and hands-on experience with at least one SQL database (Postgres, MySQL, etc) Exposure to client-side technologies such as JavaScript and React is a plus Good understanding of algorithms, data structures and their time and space complexity Curiosity, a positive attitude, and a drive to continue learning Excellent verbal and wri
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Product Okta’s Auth0 is an easy-to-implement authentication and authorization platform designed by developers for developers. We make access to applications safe, secure, and seamless for the more than 100 million daily logins around the world. Our modern approach to identity enables this Tier-0 global service to deliver convenience, privacy, and security so customers can focus on innovation. The Team The Enablement team is at the core of expanding Auth0's capabilities for B2B customers, enabling seamless and automated user lifecycle management at a massive scale. We build and own the critical features that enterprises rely on to connect their identity sources to Auth0, including Enterprise APIs and our powerful self-service capabilities . Our work is highly impactful, helping customers automate the creation, updating, and deactivation of users. This is a cornerstone for B2B SaaS applications that need to efficiently manage access for their own customers and partners. We work with NodeJS , TypeScript , PostgreSQL , MongoDB , and React to build these highly available and scalable services. What you’ll be doing: Help drive the architectural vision and strategy on the team to design and deliver powerful new enterprise APIs and functionality for our customers. Orchestrate and lead major technical projects across teams as necessary. Design, architect, code, and document large-scale distributed systems. Serve as a subject matter expert on building scalable, r
Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Software Engineer to join our Data Acquisition team. Responsibilities: Own and lead engineering projects in the area of data acquisition including web crawling, data ingestion, and search. Collaborate with other sub-teams, such as Data Processing, Architecture, and Scaling, to ensure smooth data flow and system operability. Work closely with the legal team to handle any compliance or data privacy-related matters. Develop and deploy highly scalable distributed systems capable of handling petabytes of data. Architect and implement algorithms for data indexing and search capabilities. Build and maintain backend services for data storage, including work with key-value databases and synchronization. Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks. Conduct and analyze experiments on data to provide insights into system performance. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in software development. Experience with large web crawlers a plus Strong expertise in large stateful distributed systems and data processing. Proficiency in Kubernetes, and Infrastructure-as-Code concepts. Willingness and enthusiasm for trying new approaches and technologies. Ability to handle multiple tasks and adapt to changing priorities. Strong communication skills, both written and verbal. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health is looking for hands-on, passionate people who want to join a high energy and growing team to make a difference in customers’ lives and who want to be on the forefront of digital innovation that aims to reinvent what a pharmacy and a health care company can be in the digital world. Currently, we are seeking a Staff Software Engineer – Search / AI who as a Senior technical leader, be responsible for driving architecture, design, and delivery of scalable, cloud-native platforms built on microservices architecture and AI capabilities. This role combines deep hands-on engineering with strategic leadership to build intelligent, distributed systems. The right candidate will be a strong analytical thinker and be able to simplify complex problems, processes or projects into component parts explore and evaluate them systematically. We love to collaborate and help each other and we want someone to share that ideology. Expectations for the Role Drive enterprise architecture and technical strategy with strong focus on microservices-based design and AI platform engineering Design and develop highly scalable microservices architectures, including APIs, domain-driven services, and event-driven systems Lead the development and integration of AI/ML solutions, including LLMs, Retrieval-Augmented Generation (RAG), and agentic frameworks Develop sc
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Principal Software Engineer Principal Software Engineer Position Overview The Principal Software Engineer is a senior individual contributor within Architecture & Technology (A&T), responsible for driving enterprise engineering strategy, architecture standards, and technology excellence across Mastercard. This role provides deep technical leadership across distributed systems, cloud-native platforms, resiliency, observability, and software engineering practices while influencing technology direction across multiple teams and domains. The Principal Engineer partners with senior engineering leaders, architects, and platform organizations to define architectural standards, establish reusable patterns, and guide critical technology decisions. Through technical expertise, thought leadership, and cross-functional influence, this role helps teams build secure, scalable, reliable, and operationally excellent solutions that align with Mastercard's long-term engineering strategy. Role • Provide technical leadership and architectural guidance across multiple engineering organizations, driving consistent adoption of engineering standards, best practices, and enterprise technology patterns. • Partner with platform CTOs, architects, and engineering leaders to evaluate technology investments, transformation initi
NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data
NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready for their most important workloads. In this role, you help design and run services that sit at the intersection of hardware, security, and large-scale distributed systems. We partner closely with security, silicon, platform, and cloud teams to bring new ideas into reliable production services that people rely on every day. We care about building systems that last, supporting each other, and creating space for learning and experimentation. If you enjoy solving complex problems, keeping services running smoothly, and collaborating with teammates from many disciplines, we would love to talk with you! What you’ll be doing: Your main focus will be on building and managing our core attestation cloud services. Day-to-day responsibilities include crafting APIs and integrations, boosting reliability, and working alongside NVIDIA teams to convert hardware trust mechanisms and standards into production-ready solutions. You will contribute significantly to shaping how customers verify that NVIDIA platforms are secure and prepared for their workloads. Crafting and evolving attestation cloud services, APIs, and SDK/CLI integration points that confirm the integrity of NVIDIA platforms across data center, AI, networking, and partner environments. Improving reliability and operational maturity through SLOs/SLIs, alerting, runbooks, incident response, and safe rollout practices. Crafting resilient service behavior that handles dependency failures, caching challenges, regional issues, customer-side resilience needs, and graceful degradation. Architecting trust-material distribution for certificate status, re
We are seeking a Staff Engineer to join our growing team to provide technical direction and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Engineer on this new team, you will be responsible for providing technical leadership to teams developing cutting edge technologies related to enabling deployment at scale of AI applications. You will take on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day. We value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We're looking to speak with candidates based in the New York City area for our hybrid or in-office working models. Position Expectations Work closely with product management, product engineering, product design peers as well as other teams within the company to define the first version and future evolution of the service Design, build and deliver well-tested core pieces of the platform in collaboration with other vested parties Contribute to shaping architecture, code reviews and development practices, developer experience as the teams and product grow Mentor fellow engineers and assume ownership and accountability of projects Qualifications Strong background in building core components for high scale compute and data distributed systems 8+ years experience of building distributed systems, and/or foundational cloud services at scale and an interest in working with Python, Go and Java Proven success in designing, writing, testing, debugging, performance tuning, possessing a strong grip on the foundational materials of computer science and maintaining distributed and/or highly concurrent software s
Get new distributed systems engineer jobs by email
Daily job updates · Unsubscribe anytime