Jobiba hiring network

Platform Deployment Management Lead Jobs

10,000 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current platform deployment management lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl

pythonlinuxai
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Forward Deployed Engineer at Anyscale, you will partner directly with our most strategic customers, including Spanish-speaking customers across Latin America and other regions, to ensure they achieve meaningful business outcomes with Ray and the Anyscale platform. Embedded within customer teams, you’ll act as a trusted advisor, aligning technical solutions with customer priorities, accelerating time-to-value, and driving adoption at scale. You’ll work across customer organizations — from technical leadership to individual contributors — to scope and deliver impactful solutions. By connecting insights from the field back to our product and engineering teams, you’ll help shape Anyscale’s roadmap and ensure we remain focused on solving our customers’ most critical challenges. In this role, you will: Work onsite with key customers to lead proof-of-value engagements, deployments, and enterprise adoption Translate business objectives into technical solutions that demonstrate clear ROI and strategic impact Build and deliver high-impact demos, reference architectures, and enablement programs tailored to customer needs Act as a trusted advisor across all levels of the organization, ensuring confidence in Anyscale and

kubernetesmachine learningai
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc

awsazuregcp
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. As a Site Reliability Engineer you will: Build self-service systems that automate managing, deploying and operating services. This includes our custom Kubernetes operators that support language model deployments. Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems. Take steps required to ensure we hit defined SLOs, including pa

awsazuregcp
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by leading the design of high-performance, scalable and reliable machine learning systems? Do you want to set technical direction and help shape the next generation of AI platforms powering advanced NLP applications? We are looking for a Lead Member of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will serve as a key point of contact for customers, leading the design of customized deployments to meet their specific needs, and mentoring engineers to raise the technical bar across the team. You may be a good fit if you have: 8+ years of engineering experience running production infrastructure at a large scale, with a track record of technical leadership Demonstrated experience leading the architecture

awsazuregcp
View job →
S
Snowflake
📍 Dublin• Full-time
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. At Snowflake, our People team is the backbone of an extraordinary employee experience. As a People Systems Analyst on the People Technology team, you'll serve as a key business systems partner — bridging the gap between our People Operations stakeholders and the platforms that power their work. You'll lead discovery, drive requirements, own testing, and see projects through from kickoff to hypercare — ensuring our HR and talent systems scale with the speed and complexity of Snowflake. This is a hybrid role based in our Menlo Park, CA, Dublin, CA, or Bellevue, WA offices with a minimum of 3 days per week in-office attendance required. WHAT YOU'LL DO: Partner closely with People Operations, Talent Acquisition, and HR stakeholders to conduct discovery, translate business needs into clear requirements, and recommend system solutions Lead end-to-end project delivery for system implementations and enhancements — including requirements gathering, solution design, UAT planning and execution, cutover coordination, go-live deployment, and hypercare support Create and own test plans, test scripts, and defect tracking; lead structured testing cycles and coordinate cross-functional testers to ensure quality outcomes before every release Administer, configure, and support HR systems day-

M
Mongodb
📍 United States• Full-time• From $101K/yr
1mo ago

We are seeking a Salesforce Engineer III to design, build, and support secure, scalable Salesforce solutions that support critical business operations. This is a senior individual contributor role requiring strong hands-on technical expertise, sound engineering judgment, and experience working in regulated environments. This position supports systems and data subject to FedRAMP and other U.S. government compliance requirements. Due to the nature of the work, U.S. citizenship is required. This role will be based remotely in the United States. What You’ll Do Design, build, and maintain Salesforce solutions across sales, service, customer, and operational workflows Implement functionality using Salesforce configuration, Flows, Apex, Lightning Web Components (LWC), SOQL, and platform automation Develop and support integrations between Salesforce and external enterprise systems Translate business requirements into secure, scalable, and maintainable technical solutions Participate in technical design reviews, sprint planning, code reviews, testing, deployments, and production support Troubleshoot and resolve issues related to automation, integrations, data integrity, and system performance Create and maintain technical documentation, support UAT and regression testing, and assist with user enablement Ensure all solutions comply with Salesforce best practices, internal security standards, and regulatory requirements Leverage AI-powered development tools to improve productivity and analysis, while adhering to security and compliance guidelines Responsibilities Technical Delivery: Own end-to-end delivery of Salesforce features using a mix of declarative and programmatic solutions Secure Engineering: Design and implement solutions that meet FedRAMP, security, and compliance requirements, including least-privilege access and auditability Integration Design: Build and maintain secure integrations between Salesforce and downstream systems Engineering Excellence: Maintain high st

javascriptjavamongodb
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Senior Software Develope r, you will focus heavily on building, operating, and expanding our container provisioning platforms. You will be responsible for designing resilient container infrastructure and contributing to our technical migration to Amazon EKS to improve platform reliability. In this position, you will collaborate with engineering teams across Robinhood to deliver reliable platform integrations for core capabilities like networking and security. Your work will directly help our infrastructure scale efficiently while maintaining a high standard of safety and system uptime. This role is bas

awskubernetesai
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Software Developer, you will focus on building, maintaining, and scaling our container provisioning platforms. Working alongside senior engineers, you will write code to improve our infrastructure capabilities and actively participate in our technical transition to Amazon EKS. In this role, you will collaborate with teams across the organization to ensure robust platform integrations for everyday application needs like security and networking. Your efforts will directly improve system visibility, automation, and reliability across the platform. This role is based in our Toronto office(s), with in-perso

awskubernetesai
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Software Platform team accelerates developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Compute team focuses on building and operating a highly available, scalable Kubernetes-powered container platform. We ensure that our infrastructure seamlessly supports reliable application deployments, integrates core platform capabilities, and enables multi-region scalability. We are expanding our core container systems to support our next phase of technical growth! As a Senior Software Engineer , you will focus heavily on building, operating, and expanding our container provisioning platforms. You will be responsible for designing resilient container infrastructure and contributing to our technical migration to Amazon EKS to improve platform reliability. In this position, you will collaborate with engineering teams across Robinhood to deliver reliable platform integrations for core capabilities like networking and security. Your work will directly help our infrastructure scale efficiently while maintaining a high standard of safety and system uptime! This role is based in our Be

vueawskubernetes
View job →

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. L9 - Senior Business Solutions Engineer, Legal Tech Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Legal Technology team, within the BizTech organization, leads the mission to deliver innovative technology, empowering our legal function to utilize technology productively, driving connection, and scale support. This role sits within BizTech and partners directly with Airbnb's global Legal organization to accelerate the deployment of AI-powered tools, implementation of legal systems & integrations, and workflow automation across the CLO org. The Difference You Will Make: We're looking for a world-class Senior Business Systems Engineer to help redefine how Legal operates. You'll deliver fast, practical solutions using internal tools and agentic AI — owning system and tool changes end-to-end, from design through support. You'll spot what's slowing teams down, champion best practices, and safeguard data integrity across our platforms — all while staying ahead of how AI is reshaping the way we work. This isn't conventional IT or Legal Ops. You'll be embedded directly within our Legal teams, solving real challenges in real time. As a key driver of our Legal Tech strategy, you'll lead the deployment of AI agents, legal application implementations, and automation that

aigorust
View job →
O
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 United States• Full-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We're seeking talented Robotics Software Engineers to expand our robotics data collection and evaluation program. This highly technical role involves designing, implementing, and optimizing software solutions across diverse robotics hardware. You'll work closely and collaboratively with multidisciplinary teams; including software, hardware, research, and operations - to drive advancements in our robotic systems. This role is based in San Francisco, CA, and requires in-person 4 days a week. In this role, you will: Help develop and grow our data collection labs, owning the entire integration lifecycle, from identifying and sourcing new hardware to collaborating with mechanical and electrical engineers on setup, software integration, and operational deployment. Develop innovative robot control interfaces suited to a variety of morphologies, environments, and tasks. Collaborate closely with research and engineering teams to develop automation tools and machinery that facilitate the evaluation of advanced robotic policies. Lead the design and implementation of data collection, visualization, and quality control processes. You might thrive in this role if you: Have 5+ years of professional software engineering experience developing and shipping production-quality systems in robotics or hardware-integrated environments. Have extensive experience integrating and deploying industrial automation systems, off-the-shelf robotics platforms, or custom hardware into production environments. Bring hands-on experience delivering production-quality software

awsrestai
View job →
🔔

Get new platform deployment management lead jobs by email

Daily job updates · Unsubscribe anytime