About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali
Jobiba hiring network
Distributed Systems Engineer Jobs
1,301 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current distributed systems engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Search Team at Postman is responsible for enabling users to quickly find and get started with the APIs that they are looking for. Postman is growing at a rapid pace, and this manifests into an ever-increasing volume of data that users create and consume, within their teams and in the Public API Network. We focus on improving discovery and ease of consumption over this data. We are looking for a Senior Engineer with 6+ years of experience deep backend expertise on search and ETL systems and a strong product mindset, to lead core initiatives on our search platform. In this role, you'll work at the intersection of infrastructure, relevance, and developer experience—designing systems that power search across the platform. You’ll bring a bias for action, a strong backend foundation, and the curiosity to explore beyond traditional boundaries, including areas like high performance web services, high volume data pipelines, machine learning, and relevance tuning. What You'll Do Own end to end architecture and roadmap of search platform consisting of distributed indexing pipelines, storage infra and high performance web serve
When 5% of Indian households shop with us, it’s important to build data-backed, resilient systems to manage millions of orders every day. We’ve done this – with zero downtime! 😎 Sounds impossible? Well, that’s the kind of Engineering muscle that has helped Meesho become the e-commerce giant that it is today. We value speed over perfection, and see failures as opportunities to become better. We’ve taken steps to inculcate a strong ‘Founder’s Mindset’ across our engineering teams, making us grow and move fast. We place special emphasis on the continuous growth of each team member - and we do this with regular 1-1s and open communication. Tech Culture We have a unique tech culture where engineers are seen as problem solvers. The engineering org is divided into multiple pods and each pod is aligned to a particular business theme. It is a culture driven by logical debates & arguments rather than authority. At Meesho, you get to solve hard technical problems at scale as well as have a significant impact on the lives of millions of entrepreneurs. You are expected to contribute to the Solutioning of product problems as well as challenge existing solutions. Meesho’s user base has grown 4x in the last 1 year and we have more than 50 million downloads of our app. Here are a few projects we have completed last year to scale oursystems for this growth: ● We have developed API gateway aggregators using frameworks like Hystrix and spring-cloud-gateway for circuit breaking and parallel processing. ● Our serving microservices handle more than 15K RPS on normal days and during saledays this can go to 30K RPS. Being a consumer app, these systems have SLAs of ~10ms ● Our distributed scheduler tracks more than 50 million shipments periodically fromdifferent partners and does async processing involving RDBMS. ● We use an in-house video streaming platform to support a wide variety of devices and networks.
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Mechanical Engineer is responsible for end-to-end hardware ownership of components and sub-systems for the KARNO generator, taking designs from CAD through prototype, test, and validation. Working across mechanical, electrical, software, and performance teams, this role designs and troubleshoots complex thermal and mechanical systems that must perform reliably across extreme operating environments and a wide range of fuels. The position exists to advance the development of Hyliion's fuel-agnostic power generation technology through hands-on, test-driven engineering and disciplined design execution. Duties and Responsibilities Own hardware components and sub-systems end-to-end—from concept through durability, manufacturability, serviceability, cost, weight, and validation—taking designs from CAD to hardware running on a test stand. Design components and sub-systems that must survive extreme thermal environments, perform across a wide range of fuels (20+), and push the boundaries of metal additive manufacturing. Create 3D models in NX and generate 2D prints with full GD&T per ASME Y14.5. Perform design checking and print review to ensure tolerances, processes, and material specifications align with Hyliion's GD&T standards (ASME Y14.5). Conduct fluid and thermal systems design and optimization. Perform structural and thermal FEA (ANSYS or equivalent). Install, calibrate, and read instrumentation for pressure, temperature, flow, strain, and acceleration in lab environments. Execute prototype build, test, and validation cycles early and often to identify and resolve issues in the lab rather than the field. Collaborate cross-functionally
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Manager, Electrical Engineering is responsible for the electrical systems of the KARNO generator, including high-voltage power electronics, battery systems, low- and high-voltage architecture, wiring harnesses, and the hardware that converts linear motion into electrical output. This is a working manager role: the position leads and develops a team of electrical engineers while remaining directly involved in technical execution, including circuit architecture, schematic review, and hardware bring-up in the lab. The Manager is accountable for the technical excellence, safety, and reliability of the electrical engineering function, and for establishing the design standards and review practices the team works to. The position plans team capacity, owns hiring and development for the electrical engineering staff, and partners with mechanical, controls, supply chain, and program management on system integration. KARNO systems are deployed in data center, military, and industrial applications. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Own critical electrical designs personally, including regular time at the bench and in the test cell, while leading the team as a practicing engineer. Lead the electrical engineering team in the design and development of KARNO generator electrical systems, including high-voltage power electronics, battery systems, linear generator power stages, and low-voltage controls hardwar
Role: Application Reliability Engineer Location: Gurgaon Who we are Graviton Research Capital is a privately funded quantitative trading firm striving for excellence in financial markets research. We trade across a multitude of asset classes and trading venues using a diverse range of concepts, from time series analysis and stochastic models to machine learning and statistical inference. We analyse terabytes of data to identify pricing anomalies and drive innovation in financial markets. Key Responsibilities and Deliverables The ideal candidate will possess a strong background in technical support, with a passion for problem-solving and a commitment to excellence. As an Application Reliability Engineer, you will be responsible for: Monitor production services and respond quickly to alerts, incidents, and outages to ensure smooth operation and minimal downtime. Monitor trading systems and infrastructure., Triage issues across trading support services, databases, and infra; escalate and coordinate with the right owners, and drive root-cause analysis and ensure fixes are implemented for long-term stability. Serve as the first line of defense for trading operations. Proactively identify, address recurring issues, and build automation to reduce manual intervention. Improve observability by enhancing monitoring, logging, and alerting systems. Develop and maintain operational runbooks and SLO/SLA metrics. Eligibility and Required Skills Possess a degree in a highly analytical field, such as Engineering, or Computer Science 2-5 years of experience in Python, Shell/Bash scripting. Experience with Linux and shell/bash online tools. Hands-on experience with databases (SQL, NoSQL) Strong problem-solving and analytical skills Excellent communication skills Ability to remain calm and analytical under production pressure Good to have: Familiarity with monitoring/alerting stacks (Prometheus, Grafana, ELK, etc.) Familiarity with distributed messaging (Kafka) and caching systems (Red
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus
About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and fi nancial reporting. Team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Sta ff Software Engineer,Data to be a technical lead and help architect and scale our data reliability, data infrastructure, automation and tools to meet growing business needs. You’re excited about this opportunity because you will... Own critical data systems that support multiple products/teams Develop, implement and enforce best practices for data infrastructure and automation Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Improve the reliability and scalability of our Ingestion, data processing, ETLs, Reporting tools and data ecosystem services Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We’re excited about you because... 8+ years of professional experience as a hands-on engineer and technical leader leading multiple projects 6+ years experience working in data platform and data engineering or a similar role You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software Pro fi ciency in programming languages such as Python/Kotlin/Scala 4+ years of experience in ETL orchestration and work fl ow management tools like Air fl ow Expert in database fundamentals, SQL, data reliability practices and distributed computing 4+ years of experience with the Distributed data/similar ecosystem (Spark, Presto) and streaming technologies such as Kaa/Flink/Spark Streaming Excellent communication skills and experience working
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 The Mission The Foundry is ClickUp's internal AI innovation lab — embedded inside GTM Systems and accountable for turning AI capabilities into production-grade, internally deployed products that make every GTM function faster and smarter. We build the infrastructure that powers AI-first work across Sales, Marketing, Post-Sales, and Revenue Operations. As the Senior Software Engineer on this team you will own the technical delivery of our MCP server platform, agent orchestration layer, and internal tooling — shipping production systems used daily by hundreds of ClickUp employees, and scaling your own throughput by treating AI tools as first-class engineering collaborators. What You'll Own MCP Server Platform Design, build, and operate Model Context Protocol servers that expose CRM, ticketing, analytics, and communication data to AI agents across the GTM stack Implement Okta PKCE authentication flows and RBAC policy enforcement so agents access only the data they're authorized to touch Maintain deployment infrastructure on AWS (Bedrock, Lambda, ECS, API Gateway) and contribute to GCP workloads where applicable Own observability: structured logging, distributed tracing, latency SLOs, and on-call runbooks for every production server Agent Orchestration & AI-Native Products Build and maintain multi-step autonomous agents that execute end-to-end GTM workflows — lead qualification, deal room assembly, onboarding automation, support triage, and more Architect prompt engineering frameworks, tool-call schemas, and agent evaluation harnesses that make AI behavior predictable and auditable Integrate with LLM p
At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. We’re looking for experienced engineers who have shipped applied AI systems to production and want to define what the agent-native future looks like. We are building intelligence into the core of Linear, enabling the product to orchestrate coding, proactively move work forward, and power-up every software team. You’ll work closely with product and design to transform foundation models into structured, reliable workflows embedded deeply in the core of Linear. We care deeply about keeping Linear fast, intuitive, and opinionated—AI is no exception. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the North America. You can work from anywhere within this region. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Build AI-powered product features that feel native, fast, and delightful to use Work with product and design to prototype and iterate on intelligent workflows and user interactions Design backend services to power natural language interfaces, smart suggestions, agentic workloads, and more Optimize prompts, fine-tune model behavior, and evaluate performance Help to guide our agent platform, allowing third parties to bring agents into the core Linear experience
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! About the Team The NR Lens team builds New Relic's Federated Data Platform — a distributed SQL query engine that lets customers connect external data sources (Snowflake, PostgreSQL, Google Sheets, AWS CloudWatch) and query them directly from New Relic. You'll work on the distributed query execution layer, connector architecture, and API gateway that powers cross-source JOINs and unified analytics across customer data stores. What You'll Do Design, build, and maintain cloud-native Java microservices in the NR Lens query path: SQL Gateway, Query Gateway, and data source connector plugins Develop and harden JDBC connector integrations — including connection lifecycle management, credential handling, query pushdown optimization, and security validation Improve query reliability and performance across a multi-tenant distributed SQL deployment serving 50+ customers Build and operate services on AWS (EKS, IAM/STS, S3) with infrastructure-as-code practices Implement
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads & Discovery business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver more value to our users and our advertisers. As a Machine Learning Infrastructure Engineer, you’ll build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You will: You will co-design models and systems, working at the intersection of model architecture and ML infrastructure, partnering closely with core modelers, data and AI infrastructure engineers, and product teams to push the boundaries of large-scale training and serving. Your work will span recommendation, search, and agentic applications, including large transformer architectures, LLMs, generative rankers, and efficient offline and online content-understanding systems. You will investigate model, data, and systems tradeoffs end to end—from data pipelines and distributed training to low-latency inference and production serving. This includes designing efficient KV-cache strategies, applying p
The eBPF APM team is building a zero-instrumentation observability solution that automatically discovers services on every host, supports both plaintext and TLS-encrypted traffic, classifies Layer 7 protocols, decodes service-level traffic, and reports RED (requests, errors, duration) metrics. Leveraging deep expertise in eBPF, the team operates across a wide range of Linux kernel versions, distributions, and complex customer environments. In addition to low-level networking, the team solves challenges related to protocol versioning, TLS detection across diverse languages and runtimes, and resilient performance in production systems We’re looking for a senior engineer with strong systems-level thinking and a good understanding of Linux. You should be comfortable working close to the kernel, ideally with experience in eBPF, or with a strong desire to dive into it. Proficiency in C/C++/ Go is essential, and familiarity with networking protocols, TLS internals, or distributed tracing is a strong advantage. You’ll join a high-impact team tackling ambitious technical challenges—like decoding traffic across multiple protocols, and ensuring high-fidelity metrics in complex, real-world environments. You’ll be expected to lead design and implementation efforts, contribute to roadmap planning, and collaborate across teams to ensure our solution remains robust, scalable, and frictionless for our users. This role is a great fit for engineers who thrive on low-level, performance-sensitive problems, and want to shape the future of observability through cutting-edge kernel technology. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design and build core components of our zero-instrumentation APM product using eBPF and Go Develop systems to aut
The worldwide data management software market is massive (IDC forecasts it to be $138 billion by 2026). At MongoDB, we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the center of innovation and creativity. MongoDB is seeking a Sr. Staff Software Engineer to join the Atlas Core Data Services organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product, along with the API Platform and Developer Tools. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Core Data Services organization builds the software that manages the Atlas cluster infrastructure hosted on the three major cloud providers (AWS, Azure, and GCP), as well as the software that manages the MongoDB database hosted on that infrastructure. We are constantly challenged to design features that ensure Atlas clusters are secure, available, durable, and performant while running large-scale, critical workloads. The Sr. Staff Engineer in this role will drive innovation across the organization and the company, setting technical standards and direction that enable future growth and velocity. We are looking for engineers with the experience and high standards needed to lead at that scale. Our organization champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies systems expertise to build the foundational infrastructure of a popular database, join us. Let's build a faster, more reliable, and highly scalable database platform together. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Define standards and vision for the mission-critical Atlas SaaS data
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help design and own the building, deploying and optimizing the streaming infras
Get new distributed systems engineer jobs by email
Daily job updates · Unsubscribe anytime