ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Cloud Platform Engineer, you'll envision and build robust systems and processes that ensure our infrastructure is scalable, reliable, and efficient. This can range from automating deployments and monitoring systems to optimizing performance and managing incidents. We all work closely with our users, learning from their past struggles in operationalizing ML, onboarding them onto our platform, and turning our learnings into ideas for improving Baseten. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Build and maintain scalable infrastructure to support the deployment and operation of machine learning models. Establish standards and best practices for reliability and performance across the infrastructure. Automate processes when relevant, particularly for managing CI/CD pipelines. Own products and projects end-to-end, functioning as both an engineer and a project manager, with a focus on user empathy, project specification, and end-to-end execution. Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions. Mentor junior team members and contribute to knowledge sharing within the organization. Navigate ambiguity and exercise good judgment on tradeoffs and
Jobiba hiring network
Reliability Engineer Jobs
2,027 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At Modal, we sell cloud services atop which our customers run their critical production systems. As a rapidly growing new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform, customer base, and our team. This role is for people who are deep systems thinkers, love stacking nines, and thrive from making others move faster at scale. Responsibilities include: Identifying architectural changes to improve reliability and performance. Fostering a culture of reliability across Modal’s engineering organization. Defining and implementing operational processes such as deployments, upgrades, etc. Operating systems like Kubernetes, Postgres, Redis, etc. Participating in on-call rotations, and responding to production incidents. Requirements: 5+ years of experience writing high-quality production code. 2+ years of
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers. Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform. FDE at Baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements. You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Leadership & Team Management Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional deve
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens
Anyscale Platform Engineering Leader About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud using Ray - the popular open source platform used by companies like Netflix, Uber, Instacart and others - seamless. In this position, you will guide the vision, technical direction of the team, and recruit, enable a high-performing engineering team that delivers critical values to developers and Anyscale customers by solving complex distributed systems challenges. You will oversee and drive the strategy and execution of components which includes cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components. You will closely work with our customers and our field engineering team to solve their problems, understand their challenges and make sure they are successful. We'd love to hear from you if you have: Solid engineering management experience leading produ
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. We’re looking for an Engineer to join the ML Platform team at Synthesia. Our team builds and operates the systems that allow researchers and product teams to train, serve, and deploy generative models reliably and efficiently . This includes research infrastructure, production serving systems, internal tooling, and the platform interfaces that connect them. A growing part of our mission is making these systems more automation-friendly and agent-oriented , so that workflows can increasingly be operated through reliable tooling rather than manual effort. We’re looking for a strong generalist with a systems mindset: someone who is comfortable working across infrastructure, backend systems, and tooling, and who has seen ML systems in practice. this is not a pure ML Engineer role. We’re especially interested in people who think deeply about reliability, scalability, performance, and resource efficiency in complex production environments. This is a hands-on IC role with significant ownership. You’ll help shape how our ML platform evolves as we scale the number of models, workloads, tools and teams relying on it. What you’ll do Design and improve the platform systems that support model training, evaluation, and product
As a Senior Backend Engineer on Coder's Enterprise Experience team, you'll build the systems that help large organizations run Coder in production with confidence. You'll improve how Coder scales, how it's upgraded, and how reliably it performs in regulated, air-gapped, and enterprise environments. You'll work on a genuinely cross-functional team of backend, platform, and QA engineers who own the end-to-end experience for Coder operators. From designing new features to evolving Coder's architecture, you'll partner across engineering and product to solve complex problems and ship software that operators trust. What you'll do here Design and build new features end to end, from technical design through production rollout. Design and implement backend architecture changes that support Coder's long-term scalability goals. Investigate and resolve scalability bottlenecks under production-like load, from database access patterns to concurrency handling in coderd. Improve database migration safety and upgrade reliability through schema compatibility, background migrations, and safe rollback strategies. Own the backend side of issues surfaced by Coder operators and administrators. Document the design, implementation, and operational tradeoffs of the systems you build. Participate in code reviews, RFC-style design discussions, and on-call rotations for the services you own. What we're looking for 5+ years of professional software engineering experience, including significant production experience with Go. Deep understanding of Go's concurrency model, including goroutines, channels, the sync package, and debugging race conditions under real-world load. Experience designing and operating relational databases in production, including schema design, migrations, and transactions. Strong verbal and written communication skills. Exceptional debugging and troubleshooting skills, with the persistence to drive complex problems to resolution. A self-motivated, analytical engineer who enj
As an Engineering Manager on Coder’s Agentic Engineering team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction, grow the team, and keep execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Agentic Engineering organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development enviro
Senior Software Engineer - Analytics Compute Platform Team About the Role & Team Every chart, insight, and experiment result a customer sees in Amplitude passes through the compute layer that this team owns. Our team sits between Amplitude's front-end analytics products/ rest APIs/ MCPs and Nova, our proprietary analytics database, and owns the compute API that translates a user's question into a fast, correct answer. Our mandate: enable a highly performant, reliable, and flexible way to compute insights. Those three goals are often in tension the more flexible we make the system, the harder it is to keep it fast and simple, and a lot of the interesting engineering work on this team lives in that tradeoff. Day to day, you'll work closely with engineers on data management, marketing analytics, Session Replay, Experiment, CDP, Query, and various product teams, since they're all consumers of what we build. As a Senior Software Engineer, you will Design and build core parts of the query/compute engine that powers analysis across Amplitude's product suite Evolve the compute API and semantic layer that other engineering teams build features on top of, so that it can support a more generic table Help unify and evolve our core data model, including how non-event data (profile properties, lookup tables, etc.) is represented alongside event data Improve the performance, reliability, and scalability of query planning and execution on the Amplitude query engine Partner with teams across Session Replay, Experiment, CDP, Query, and frontend to understand their needs and shape the compute layer around them Take ownership of projects end-to-end, from design through rollout You'll be a great addition to the team if you have Strong backend or distributed-systems engineering experience, ideally touching OLAP databases, query engines, or data infrastructure Experience with SQL, query planning/execution, or building APIs that many other engineering teams depend on Comfort reasoning
About the Role & Team We’re looking for an Engineering Manager to lead the Delivery and SDK team within Statsig at Amplitude. This team builds the systems that connect developers to Statsig and safely deliver product experiences to end users around the world. The team owns two foundational areas: Delivery: The globally distributed systems that evaluate Experiments and Feature Gates and deliver the resulting experiences to end users. These systems sit directly in the critical path of our customers’ applications, making availability, scale, correctness, and latency essential. SDKs: More than 30 SDKs spanning web, mobile, server, edge, TV, gaming, and emerging platforms. Our goal is to meet developers wherever they build and give them an ergonomic, reliable way to integrate Statsig into any application or technology stack. This is a uniquely broad technical leadership role. You’ll work across distributed systems, mobile platforms, server runtimes, edge environments, and developer tooling—sometimes all in the same week. We’re looking for a hands-on, technically versatile leader who enjoys moving between technologies, learning unfamiliar systems, and solving problems across abstraction boundaries. You don’t need to be the world’s foremost expert in a single language or platform. You do need the curiosity and technical judgment to work effectively across many of them—and the ability to build a team that can make a complex, polyglot ecosystem feel simple and dependable to developers. What You’ll Do Lead and grow the team responsible for Statsig’s global delivery infrastructure and ecosystem of 30+ SDKs. Experience leading a team of 8-10 team members Define the technical strategy for highly available, low-latency evaluation and delivery systems operating at global scale. Make Statsig exceptionally easy to adopt by continuously improving SDK ergonomics, performance, reliability, consistency, and documentation. Expand our SDK coverage as new languages, frameworks, runtime
About The Role & Team The Developer Experience (DX) team at Amplitude builds and maintains the foundations that power how developers integrate, extend, and trust Amplitude across platforms. Our mission is to make Amplitude’s SDKs reliable, easy to adopt, and a joy to build on, so customers can confidently instrument their products and unlock insights at scale. We’re looking for a Senior Software Engineer, iOS to play a key technical leadership role on our DevEx team. In this role, you will lead the design and development of Amplitude’s core iOS SDKs, including Analytics and Session Replay , and serve as the iOS platform expert that other SDK teams, such as Experiment, Guides, and Surveys rely on. As a Senior Engineer, you’ll operate with a wide scope and high impact: setting technical direction for the iOS platform, driving cross-SDK architecture, improving performance and reliability, and raising the bar for developer experience across Amplitude’s mobile ecosystem. As a Senior Software Engineer, you will: Lead the technical direction, architecture, and long-term evolution of Amplitude’s iOS SDK platform. Own and drive development of core iOS SDKs, including Analytics and Session Replay, with a strong focus on performance, reliability, and ease of use. Act as the iOS platform expert and trusted partner for other SDK teams (Experiment, Guides & Surveys), enabling them to build on shared foundations safely and efficiently. Design and evolve shared infrastructure, APIs, and abstractions that scale across multiple iOS SDKs. Collaborate closely with Product, and Customer Support to ensure SDKs meet real customer needs. Lead cross-team technical discussions, reviews, and architectural decisions that span multiple SDKs. Improve developer experience through better APIs, documentation, tooling, testing strategies, and sample apps. You'll be a great addition to the team if you have: A strong focus on developer experience and empathy for the engineers who use your work
About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically. We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost-efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers. What You’ll Do Build and evolve core query engine infrastructure Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Design for high-throughput automated quer
About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We're entering a world where AI agents don't just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents' ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova's throughput, correctness, and operational rigor grows dramatically. We're looking for a Senior Software Engineer who wants to go deep on the engine internals and the infrastructure underneath. You'll own significant components of a modern OLAP system — across query execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — and drive meaningful improvements to performance, cost-efficiency, and reliability. You'll grow your technical influence through the quality of your code, your design contributions, and your collaboration with other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of enterprise customers. What You'll Do Build and improve core query engine components Contribute across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Help ensure Nova's components support
About the Role & Team We’re looking for an Engineering Manager to lead the Delivery and SDK team within Statsig at Amplitude. This team builds the systems that connect developers to Statsig and safely deliver product experiences to end users around the world. The team owns two foundational areas: Delivery: The globally distributed systems that evaluate Experiments and Feature Gates and deliver the resulting experiences to end users. These systems sit directly in the critical path of our customers’ applications, making availability, scale, correctness, and latency essential. SDKs: More than 30 SDKs spanning web, mobile, server, edge, TV, gaming, and emerging platforms. Our goal is to meet developers wherever they build and give them an ergonomic, reliable way to integrate Statsig into any application or technology stack. This is a uniquely broad technical leadership role. You’ll work across distributed systems, mobile platforms, server runtimes, edge environments, and developer tooling—sometimes all in the same week. We’re looking for a hands-on, technically versatile leader who enjoys moving between technologies, learning unfamiliar systems, and solving problems across abstraction boundaries. You don’t need to be the world’s foremost expert in a single language or platform. You do need the curiosity and technical judgment to work effectively across many of them—and the ability to build a team that can make a complex, polyglot ecosystem feel simple and dependable to developers. What You’ll Do Lead and grow the team responsible for Statsig’s global delivery infrastructure and ecosystem of 30+ SDKs. Leading a team of 8-10 team members Define the technical strategy for highly available, low-latency evaluation and delivery systems operating at global scale. Make Statsig exceptionally easy to adopt by continuously improving SDK ergonomics, performance, reliability, consistency, and documentation. Expand our SDK coverage as new languages, frameworks, runtimes, and comp
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime