About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're building a platform that covers the whole life of an LLM: training it, deploying it, and observing it in production. We already run multi-node training, elastic inference, sandboxes, and distributed volumes, and we control the infrastructure underneath. We’re looking for research depth in post-training to sit alongside our systems and product work. What you'll do: We are looking for research scientists with a strong track record in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This role is well suited to candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and improving inference-time efficiency, reliability, and robustnes
Jobiba hiring network
Platform Engineering Lead Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current platform engineering lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Staff Software Engineer on the Staking Platform team within the Platform group, you'll serve as Coinbase's definitive Solana staking technical authority, owning strategy across validator operations, staking integrations, and protocol evolution. Coinbase operates one of the largest Solana staking operations in the world, managing approximately 9.25% of all staked SOL. You'll combine deep Solana protocol mastery with hands-on engineering execution and external ecosystem influence to shape Coinbase's Solana staking trajectory. What you'll do: Own Coinbase's multi-year technical strategy for Solana staking across validator performance, protocol participation, and product integration, connecting engineering decisions to yield optimization, cost efficiency, and customer growth. Lead the engineering effort to achieve industry-leading APY through validator optimization, including vote accuracy, block production, MEV strategies, commission tuning, and stake distribution tooling. Serve as Coinbase's foremost authority on the Solana runtime, consensus mechanism, staking economics, and validator client landscape (Agave, Firedancer), evaluating protocol upgrades (e.g., SIMD proposals) and proactively positioning Coinbase for changes before they land. Partner with Retail and Institutional Staking product and engineering teams to architect scalable staking integrations across
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll make cold starts feel local when the data is hundreds of milliseconds away, designing the caching, preloading, and peer-to-peer layers that hide object-store latency and keep public ingress off saturated uplinks. You'll own durability and cost at petabyte scale, from streaming and batch replication between origins, to garbage collecti
Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Job Description: ABOUT THE TEAM Our Engineering team is revolutionizing how Appian resolves customer issues by building AI-powered diagnostic tools and internal platforms that dramatically reduce resolution time. We combine deep technical expertise with AI innovation to empower support engineers to identify root causes faster, ensuring our customers experience uninterrupted business-critical operations. This team operates with high autonomy and directly impacts customer satisfaction and retention across Appian's enterprise client base. THE OPPORTUNITY Appian is at an inflection point where intelligent automation meets customer success, and we need a technical leader who can architect and ship the next generation of diagnostic platforms. This role exists because our growing enterprise customer base demands faster, smarter issue resolution, and AI agents are the key to unlocking that scale. You will shape both the technical direction and the team culture while delivering tools that become essential to how Appian supports thousands of mission-critical deployments worldwide. WHAT YOU'LL DO Lead the design and delivery of a scalable AI-powered triage platform that reduces mean time to resolution by leveraging AI agents, LangChain frameworks, and intelligent diagnostics Lead the engineering team building AI-powered diagnostic tools that cut customer issue resolution time in half while scaling a platform that empowers support engineers across the globe. Architect quarterly releases of diagnostic tools on AWS infrastructure, balancing innovation velocity with pro
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed Engineers on our engineering team who want to work at the intersection of deep infrastructure work and direct customer impact. As an FDE, you'll partner with leading AI companies and foundation labs on cloud architecture, networking, storage, containerization, sandboxing, and more — helping them design and ship production infrastructure on Modal's platform. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the infrastructure stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and deploy massive-scale production workloads on Modal Lead technical discovery and architect
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Modal is seeking an experienced Forward Deployed Engineer (FDE) to partner with our sales team and drive technical sales success. As an FDE, you will be the technical voice in our sales process, working directly with Account Executives to help enterprise customers understand how Modal can transform their AI/ML infrastructure. You will: Partner with Account Executives to identify, qualify, and close strategic enterprise opportunities Lead technical discovery sessions with prospective customers to understand their current infrastructure, pain points, and requirements Design and present compelling technical solutions that demonstrate how Modal addresses customer needs Architect migration paths from existing cloud infrastructure (AWS, GCP, Azure) to Modal's serverless platform Conduct technical demos, experiments, and proof-of-concepts that showcase Modal's capabilitie
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our NYC or SF office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our Stockholm office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We’re looking for an Infrastructure Security Engineer to design and secure the core systems that power our platform. This role focuses on building security directly into our infrastructure—from container isolation and orchestration to identity and secrets management in a multi-tenant, cloud-native environment. You’ll work closely with engineering teams to define secure primitives and ensure our platform is resilient, scalable, and trustworthy by design. This is a hands-on, deeply technical role focused on real systems, not compliance or policy. What You'll Do: Platform & Runtime Security Design and improve isolation mechanisms for multi-tenant workloads (containers, sandboxing, execution environments) Strengthen boundaries between customers, workloads, and internal systems Identify and mitigate risks in distributed, dynamic compute environments Container &
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for engineers with deep AI/ML and low-level systems experience who want to build the best technical support experience in the world. This isn't a traditional support role — it's an engineering role where you happen to be closest to our customers. You'll split your time roughly 50/50 between working directly with customers and shipping fixes, features, and automation that improve Modal for everyone. When you help a customer debug a training run, you'll also fix the underlying issue in the platform. When you notice ten customers hitting the same friction point, you'll build the tooling or automation that eliminates it entirely. This role is for people who solve problems, not people who answer tickets. The problems you encounter are deeply technical and arise from running some of the most demanding AI workloads in the world. You'll be a member of our eng
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Staff engineer (L4), Twilio’s Segment team. About the job As a Staff Engineer on the Twilio Segment Data platform/ pipelines team, you’ll build and scale systems that process several hundred thousands of data points per second. You will lead the development of high-scale ingestion and data processing systems You'll be designing, operating and maintaining complex distributed systems, ensuring reliability, performance, and cost-efficiency while querying petabytes of data for our customer data platform (CDP). Responsibilities In this role, you’ll: Design and deliver robust, high-scale routing experiences for the Data platform/ pipelines team for Twilio Segment. Ship features that opt for high availability and throughput with eventual consistency Collaborate with engineering and product leads, as well as teams across Twilio Segment Support the reliability and security of the platform Build and optimize globally available and high
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our New York City, Austin, Seattle or San Francisco offices, or work fully remotely on standard East Coast business hours. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our Dublin office, or work fully remotely in Ireland. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations and processes A deep understanding of Linux and networking concepts,
Get new platform engineering lead jobs by email
Daily job updates · Unsubscribe anytime