We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations. What You Will Be Doing: Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale. Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams. Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals. Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation. Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams. Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness. Provide technical leadership and mentorship, influence engineerin
Role overview
Job description
We're looking for a Senior Platform Reliability Engineer who brings strong software engineering skills and a deep understanding of system behavior under load and stress. This role is a good fit for someone who wants to own reliability as a first-class concern – building the foundational systems that protect Asana's platform, not just responding when things go wrong. You'll build core platform systems like load shedding, rate limiting, circuit breakers, and traffic controls that protect Asana under real-world load. This is deep, cross-cutting work that shapes stability and performance of our entire infrastructure – and you'll partner closely with other platform teams to make reliability something that's built in, not bolted on. Our tech stack includes: AWS, Kubernetes (EKS), CloudFront, Istio, Cilium, MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, an
…What they are looking for
Skills & requirements
Qualification
7+ years of experience building and operating backend systems at scale
Department · Infrastructure Engineering
Hiring company
Asana
We are looking for enthusiastic collaborators who are passionate about their craft to be a part of our journey building technology that is a force for positive change in the world. Make an impact by helping us achieve a powerful mission while developing your career.
Keep exploring
Similar active roles
Fresh roles matched to this title and market.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Our team builds and operates Snowflake's Streaming Platform — the core infrastructure responsible for bringing data into Snowflake continuously and at scale. This includes Snowpipe Streaming and Datastream, services used by large set of enterprise customers to move mission-critical data in real time. Data is the fuel powering the new enterprise AI world, and we are the team responsible for getting it there reliably, continuously, and fast. Responsibilities Design, build, and maintain core components of Snowflake's streaming ingestion platform Improve service reliability, scalability, and latency under high-throughput production workloads Debug and resolve incidents in a distributed, multi-tenant cloud service Contribute to the design and evolution of streaming APIs, SDKs, and server-side protocols Write comprehensive tests including unit, integration, and chaos/fault-injection scenarios Collaborate with partner teams (storage, query, compute) on cross-cutting platform concerns Participate in code reviews, on-call rotations, and architecture discussions Required Qualifications 3–5 years of software engineering experience on large-scale distributed systems or cloud services Strong proficiency in Java or C++ Deep understanding of distributed systems concepts: consistency, faul
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Everpure Cloud Azure Native is a generally available service that brings enterprise-grade block storage natively to Azure. With the first version already live, we are expanding the service into new use cases and markets while improving its reliability, operability, and customer experience. Our Foundation team owns a Go-based control-plane service that coordinates how customers provision and use storage in Azure. Around it, we work with a modern cloud stack including Temporal and other platform services for workflows, automation, and observability. This is a production cloud service: the code you write directly shapes how customers deploy, scale, and operate storage in their Azure environments. You’ll work on a core storage service in a major public cloud , as part of a joint effort between Everpure and Microsoft. You’ll design and evolve APIs and service behavior in the critical path of real customer workloads, collaborating closely with engineers across both companies. Clear API contracts, long-lived interfaces, test automation, and CI/CD are fundamental to how we build. You’ll have the opportunity to own services end to end and solve complex distributed-systems problems in the public cloud. WHAT YOU'LL DO Own and evolve a production cloud service that powers Everpure Cloud Azure Native, taking features from idea and design through deployment and operation in Azure for real customers. Build new capabilities and i
$118.7K – $160K/yr
Healthcare is complex. We’re here to change that. RVO Health is a health technology company on a mission to make health easier to navigate, more accessible, and more affordable for everyone. Here, you'll help over 40 million people every month, with a team that genuinely cares about the work and each other. AT A GLANCE The Senior Software Engineer is a crucial role within our organization, requiring work in various capacities and adaptation to different work arrangements based on the needs set by the business. The successful candidate will be responsible for fulfilling their job duties in the following work situations: Where You'll Be Location: Denver, CO | Hybrid We believe great collaboration happens when we're together, solving problems, learning from each other, and connecting as a team. That's why we’re in our offices Tuesday through Thursday each week. You are welcome to work remotely Mondays and Fridays if you wish. Address: 1801 California St. Denver, CO 80202 What You’ll Do Lead the end-to-end design, development, and implementation of sophisticated software applications and systems aligned with business goals. Collaborate closely with stakeholders including product managers, designers, and other engineers to gather requirements and translate them into robust technical designs and solutions. Write high-quality, efficient, maintainable, and scalable code adhering to best practices and company standards. Debug, analyze, and resolve complex software defects and performance bottlenecks to ensure optimal system reliability and user experience. Conduct comprehensive testing and validation including unit, integration, and performance testing to guarantee software quality. Mentor and provide technical guidance to junior and mid-level engineers, fostering professional growth and knowledge sharing. Perform thorough code reviews to maintain high code quality, enforce coding standards, and promote best p
DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Senior Software Engineer – Platform Operations based in Pune, India, who will play a key role in ensuring the reliability, performance, and operational excellence of DeepIntent's platform and data ecosystem. This role requires a strong engineering mindset with the ability to troubleshoot complex technical issues, understand distributed data architectures, and collaborate across Engineering, Product, Analytics, and Customer-facing teams to deliver timely and effective solutions. As part of the Operations organization, you will work closely with Engineering to support production systems, improve operational processes, and drive platform stability. The ideal candidate is a self-motivated problem solver who is passionate about learning new technologies, improving system reliability, and delivering exceptional customer outcomes through engineering excellence. Serve as the engineering interface between Customer-facing teams, Analytics, Product, and Engineering organizations. Partner with Platform Support, Client Success, and other customer-facing teams to investigate and resolve complex platform-related issues. Analyze application, API, and data pipeline issues to identify root causes and drive timely resolution. Develop and standardize operational tools, and interfaces to support analytical and operational use cases. Monitor data pipeline executions, investigate failures, and implement corrective and pre
🔔 Get job alerts
New Senior Software Engineer, Platform Reliability jobs in Warsaw, straight to your inbox.
No spam · Unsubscribe anytime