A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed Infrastructure Engineers (FDIEs) build, operate, and maintain the infrastructure that powers Palantir’s platforms and production deployments. As an FDIE intern, you’ll work alongside full-time FDIEs to deploy and operate Palantir software across real production environments, automate manual processes, and develop novel solutions to infrastructure challenges using tools like Foundry and Apollo. Every day looks different — you might be debugging a distributed systems issue, building automation to replace a manual runbook, or designing infrastructure improvements that scale across multiple deployments. You’ll be treated as a full member of the team, with real ownership over the work you take on. Core Responsibilities As an FDIE intern, your responsibilities look similar to those at a small startup, with the resources, stability, and mentorship of an established tech company. You’ll work in small teams with minimal supervision and own end-to-end execution of real infrastructure projects. Your day might span discussing systems architecture with fellow engineers, debugging a production issue, building automation to eliminate a manual process, or deploying new Palantir products across production environments. FDIE interns are treated just like full-time engineers, with significant freedom and ownership over their work. Specifically, you can expect to: Deploy and operate Palantir software across production environments, including monitoring, alerting, configuration management, and upgrades Debug, improve, and optimize Palantir’s services and infra
Jobiba hiring network
Platform Deployment Management Lead Jobs
9,875 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current platform deployment management lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Site Reliability Operations Analyst you are the engine behind Palantir deployments. You are responsible for crafting, implementing and executing processes to streamline workflows and reduce friction. You track and stabilize projects, remove roadblocks, and anticipate customer needs to free up our engineers to focus their time and attention on the technical problems they are best equipped to solve. This position requires a combination of project management, process optimization, and execution skills. You are a person who loves fixing problems and always embraces the best idea, even when it is not your own.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Site Reliability Operations Analyst you are the engine behind Palantir deployments. You are responsible for crafting, implementing and executing processes to streamline workflows and reduce friction. You track and stabilize projects, remove roadblocks, and anticipate customer needs to free up our engineers to focus their time and attention on the technical problems they are best equipped to solve. This position requires a combination of project management, process optimization, and execution skills. You are a person who loves fixing problems and always embraces the best idea, even when it is not your own.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Site Reliability Operations Analyst you are the engine behind Palantir deployments. You are responsible for crafting, implementing and executing processes to streamline workflows and reduce friction. You track and stabilize projects, remove roadblocks, and anticipate customer needs to free up our engineers to focus their time and attention on the technical problems they are best equipped to solve. This position requires a combination of project management, process optimization, and execution skills. You are a person who loves fixing problems and always embraces the best idea, even when it is not your own.
Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Customer Engineering organization serves as the primary technical interface for our global customers and helps ensure that Micron solutions are seamlessly integrated into next-generation technology platforms. Our team coordinates deep engineering engagements, technical product qualifications, and architecture reviews to secure design wins and grow market share. We are dedicated to providing world-class technical support that accelerates revenue and builds long-term strategic partnerships with our customers. Role Description: The Field Applications Engineer - Associate plays a vital role in achieving the technical achievements needed to launch products. You will be an important member of our field engineering team. In this position, you will oversee many end-to-end technical tasks, including sophisticated product sample deployments, lifecycle management, addressing technical questions, and other customer interactions. You should adopt an "automation-first" approach to actively find ways to improve processes and apply AI-based workflows in daily operations when possible. This role can lead to a full-time FAE career for individuals who achieve the right results. It also requires demonstrating the right skills and behaviors. Example Responsibilities: Coordinate the entire process of tech
Who We Are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Team The Core Change Management group is responsible for the systems that let every Stripe engineer ship code, configuration, and infrastructure changes safely and at high velocity. You will be embedded primarily on the Service Deployments team — the owners of Stripe's end-to-end code deployment platform — with regular collaboration with the Resource Automation and Feature Deployments teams. What Makes This Role Compelling You own the foundation of how Stripe ships software. The deployment platform sits in the critical path of every engineer's workflow at Stripe. The decisions you make affect thousands of deploys per day across hundreds of services, directly determining how fast and safely Stripe's product evolves. Technically rich, architecturally active. The team is executing several concurrent platform transformations: containerizing host-based services at scale, adding intelligent multi-service deploy pipelines, extending real-time anomaly detection to earlier stages of traffic shifts, and rebuilding deployment event infrastructure on top of a durable message bus. This is not maintenance work — the architecture is in motion. Broad surface area, real ownership. You will span the full stack from container scheduling and deployment orchestration business logic to the developer-facing internal platform UI. The problems are multi-layered: reliability, developer experience, performance, and safety all at once. Your judgment prevents incidents.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will help develop and validate next-generation server platforms that power a reliable, high-performing, and cost-efficient infrastructure at scale. You will work across platform bring-up, firmware qualification, hardware validation, fleet integration, and performance optimization to support large-scale production deployments. You Will: Bring-up & Sustaining: Drive key aspects of the hardware development lifecycle, including feasibility studies, hardware bring-up, validation, deployment, and ongoing production support. Platform Optimization: Perform platform integration, performance characterization, and system-level debugging across compute infrastructure, focusing on hardware optimization, driver tuning, and thermal/power efficiency. Hardware Validation: Develop and execute rigorous evaluation and stress-testing strategies for server platforms to ensure reliability and performance under production-scale workloads. Firmware & Fleet Enablement: Support BIOS/BMC firmware qualification, hardware health monitoring, and automation tooling for firmware deployment and lifecycle management. Vendor & Cross-Functi
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for a OrioleDB Deployment Engineer to join our OrioleDB Team and help elevate our OrioleDB offering. You’ll work closely with the OrioleDB team, playing an instrumental role in technical decision-making and refining internal methodologies. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. What You’ll Be Responsible For Package software into our supabase/postgres repo using Nix (with flakes), and help us transition our packaging from traditional to Nix packaging more over time. Manage OrioleDB release lifecycles, ensuring timely major, minor, and extension upgrades. Expand platform release systems to allow developers to increasingly self-service. Optimize CI/CD and tooling, specifically expanding GitHub Actions, team tooling, and testing/release approaches. Resolve production issues by proactively identifying and fixing problems in customer deployments. Maintain best practices and tests to ensure enhanced stability and decreased deployment risks. co-owning the integration of OrioleDB into the supabase product You Might Be a Good Fit If You Have 3+ years of experience with PostgreSQL and its ecosystem, including extensions and performance optimization. Are an Infrastructure Expert with proven experience in management, tooling, and optimization. Are proficient in the Nix package management system (including flakes) alongside Ansible, Packer, Docker, QEMU/KVM, AWS, and Kubernetes. Have experience building for multiple architectures , specifically Linux and Darwin/macOS aarch64 targets. Are comfortable with polyglot environments , including builds for C/C++, Go, JavaScript, a
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for a Postgres Deployment Engineer to join our PostgreSQL team and help elevate our PostgreSQL offerings. You’ll work closely with the PostgreSQL team, playing an instrumental role in technical decision-making and refining internal methodologies. This role is ideal for someone who thrives in async, fast-paced environments and is excited about building developer tools that scale to millions. The focus of this role is owning the stability and deployment of our products. You will act as a bridge between product and infrastructure teams to improve the reliability of deployments and upgrades. You will have co-responsibility for builds and deployments via our public supabase/postgres GitHub repository, which bundles features into Docker images and AWS AMIs for cloud and local use. What You’ll Be Responsible For Package software into our supabase/postgres repo using Nix (with flakes), and help us transition our packaging from traditional to Nix packaging more over time. Manage PostgreSQL lifecycles, ensuring timely major, minor, and extension upgrades. Expand platform release systems to allow developers to increasingly self-service. Optimize CI/CD and tooling, specifically expanding GitHub Actions, team tooling, and testing/release approaches. Resolve production issues by proactively identifying and fixing problems in customer deployments. Maintain best practices and tests to ensure enhanced stability and decreased deployment risks. You Might Be a Good Fit If You Have 3+ years of experience with PostgreSQL and its ecosystem, including extensions and performance optimization. Are an Infrastructure Expert with proven experience in management, tooling, and optimization. Are p
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking an experienced Platform Software Engineer for our Systems Software Team. You will be working as part of a dynamic team and will be responsible for designing, developing, and testing system software functionality for Pure’s upcoming platforms. The work spans the gamut of Systems software and you will have the opportunity to work in a wide range of areas and features ranging from Platform drivers to networking and storage layers. WHAT YOU'LL DO Plan and influence the lifecycle of new Hardware Platforms. Work on problems ranging from design, bring up, to deployment, upgrades and fleet level reliability. Participate in the full lifecycle of new hardware platforms from early bring up through manufacturing release. Work closely with peer teams to debug complex HW/FW of new server hardware, including CPUs, chipsets, and peripheral components. Debug complex HW/FW issues across x86, PCIe, NVMe, and networking using lab tools (oscilloscope, logic analyzer, JTAG) and kernel/driver traces. Design, implement and improve remote server management capabilities (e.g., using standards like Redfish) and enhance Reliability, Availability, and Serviceability (RAS) features. Design, write and maintain software components in C/C++, Python, Golang and RUST. Collaborate with vendors on requirements specification and follow through to system delivery. Work closely with hardware engineers, system architects, and o
JOB TITLE Access and Identity Management - Business Analyst A CAREER WITH POINT72'S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source and AI solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU'LL DO Translate business and security requirements into scalable, production‑ready AIM and IGA solutions that support firmwide operations Partner closely with business, Global Systems Administration, and technology stakeholders to deliver secure, efficient, and user‑friendly access solutions Analyze data to identify access trends, control gaps, and process inefficiencies, with a focus on automating manual processes and improving platform scalability Own requirements gathering, functional documentation, and user stories from ideation through UAT, deployment, and post‑release support Leverage tools such as SailPoint, JIRA, SQL, Splunk, and Power BI to support access governance, reporting, and operational insights Design and maintain dashboards and KPIs that measure access lifecycle performance, risk indicators, and user experience Support data‑driven decision making by validating data quality, ensuring consistency, and aligning metrics to business objectives Build deep domain expertise in identity governance, security controls, and investment‑driven technology platforms Contribute to a culture that prioritizes integrity, transparency, and the highest ethical standards WHAT’S REQUIRED 4+ years of experience as a Business Analyst supporting Access & Identity Management (AIM) and Identity Governance & Administration (IGA) initiatives
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Cloud Platform Engineer, you'll envision and build robust systems and processes that ensure our infrastructure is scalable, reliable, and efficient. This can range from automating deployments and monitoring systems to optimizing performance and managing incidents. We all work closely with our users, learning from their past struggles in operationalizing ML, onboarding them onto our platform, and turning our learnings into ideas for improving Baseten. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Build and maintain scalable infrastructure to support the deployment and operation of machine learning models. Establish standards and best practices for reliability and performance across the infrastructure. Automate processes when relevant, particularly for managing CI/CD pipelines. Own products and projects end-to-end, functioning as both an engineer and a project manager, with a focus on user empathy, project specification, and end-to-end execution. Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions. Mentor junior team members and contribute to knowledge sharing within the organization. Navigate ambiguity and exercise good judgment on tradeoffs and
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a skilled and experienced Senior AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay curre
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay current with ad
Get new platform deployment management lead jobs by email
Daily job updates · Unsubscribe anytime