- Proven experience deploying and managing Kubernetes clusters for AI/ML workloads. Experience of at scale deployments with Azure Kubernetes. Experience level - 5 Years or more Positions - 2 Proven experience deploying and managing Kubernetes clusters for AI/ML workloads. - Experience of at scale deployments with Azure Kubernetes Service, RedHat OpenShift, Microk8s and Helm Charts. - Expertise with infrastructure and resource management and virtualization tools such as VMWare/EXSi, KVM, Ansible, Redfish. - Strong understanding of Run:AI platform, including job scheduling, quota management, and GPU virtualization. - Knowledge of NVIDIA AI Enterprise components including, NIM, NeMO, TAO, Triton and Nucleus Servers - Familiarity with DGX systems, Jetson, and NVIDIA’s AI Factory components. - Proficiency in Python, C++, and optionally .NET/C# for enterprise integration.
Jobiba hiring network
Platform Deployment Management Lead Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current platform deployment management lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Defense Sector at Leidos is seeking a motivated TS/SCI cleared Network Administrator to support the installation, configuration, and day-to-day management of enterprise network infrastructure. This role is an excellent opportunity for an early-career network professional to gain hands-on experience with routing and switching platforms, including Session Smart Router (SSR) / 128 Technology SD-WAN solutions, in a structured and security-conscious environment. The ideal candidate demonstrates a solid foundation in networking fundamentals, a willingness to learn vendor-specific technologies, and the discipline to operate within DoD network standards. The job duties will be performed daily on site at Langley Air Force Base, VA. Roles and Responsibilities: Assist in the configuration, deployment, and ongoing management of routers, switches, and Session Smart Router (SSR) appliances across enterprise and edge network environments. Support the design and implementation of routing policies, service policies, and traffic steering configurations on SSR/128 Technology platforms under senior engineer guidance. Perform LAN switching administration — including VLAN configuration, spanning tree, trunking, and port security — on Juniper EX Series and/or Cisco Catalyst platforms. Assist with the configuration and troubleshooting of routing protocols including OSPF and BGP (eBGP and iBGP) across enterprise WAN and data center environments. Monitor network health, availability, latency, and throughput using network management tools; escalate anomalies and assist in root cause analysis. Support configuration and maintenance of firewall rules, access control lists (ACLs), IPsec VPN tunnels, and other network security controls. Execute software and firmware upgrades, patch management, and lifecycle maintenance activiti
NVIDIA is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you will be doing: Design and develop firmware solutions for manageability and observability of data center servers. Actively participate in hardware bring-up activities, OOB firmware development, protocol stacks (Redfish, PLDM, MCTP, NSM) and hardware-software co-design for Cloud Service Provider deployments. Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Deep expertise in embedded firmware, server management controllers, and hardware bring-up with proven track record of shipping production BMC solutions Strong knowledge of DMTF protocols (Redfish, IPMI, PLDM, MCTP, SPDM), telemetry frameworks, and out-of-band management architectures Expert-level skills in C/C&
CRM Executive Location: Noida Experience: 1–3 Years Business: Paytm Money About the Role Paytm Money is looking for a detail-oriented CRM Executive to join its Execution Pod. In this role, you will be responsible for the end-to-end production, quality assurance, and deployment of customer communications across WhatsApp and Push Notification channels. You will ensure that every customer communication is accurate, engaging, compliant with SEBI regulations, and delivered seamlessly. This role requires strong attention to detail, excellent content review skills, and hands-on experience with CRM platforms. Key Responsibilities Draft, edit, and finalize communication copy for marketing campaigns in accordance with SEBI advertising and communication guidelines. Review and proof creatives, subject lines, CTAs, and campaign content before deployment to ensure accuracy and compliance. Set up, test, QA, and schedule campaigns on CleverTap across WhatsApp and Push Notification channels. Own campaign execution processes, including checklists, approvals, stakeholder coordination, and go-live timelines. Partner closely with Compliance and Legal teams to secure mandatory approvals and disclosures. Maintain campaign calendars, trackers, and execution dashboards to ensure timely delivery of all customer journeys. Monitor campaign deployment and proactively identify audience mismatches, delivery issues, or QA gaps. Ensure operational excellence with zero compliance misses and flawless campaign execution. Desired Skills & Qualifications 1–3 years of experience in CRM, Lifecycle Marketing, Content Operations, Marketing Operations, Customer Engagement, or Campaign Management. Hands-on experience with CleverTap is strongly preferred. Experience with MoEngage, WebEngage, Netcore, Braze, or similar CRM platforms is also valuable. Strong written English and content editing skills with the ability to create crisp, customer-friendly communication. Understanding of financial services com
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. Key Responsibilities Develop, support, and maintain automated CI/CD pipelines to streamline application delivery. Standardize build, release, and deployment processes across applications and platforms. Establish repeatable, auditable, and traceable release management practices to ensure deployment consistency and compliance. Manage and support Git-based source control, branching strategies, and release workflows. Integrate automated testing, quality gates, security controls, and compliance checks into deploym
About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal
Associate and Mid-Level Software Engineers Company: The Boeing Company The Boeing Company is looking for an Associate or Mid-Level Software Engineer to join our team in Berkeley, MO. Position Responsibilities: Lab Environment Provisioning: Design, provision, and maintain scalable lab environments on-premises and on cloud platforms (Azure, AWS) using Infrastructure as Code (IaC) tools such as Terraform and Ansible. Automated Integration Testing Orchestration: Develop and manage automated integration testing workflows that aggregate inputs from multiple program segments, ensuring comprehensive test coverage and timely feedback. Deployment Management: Manage and optimize automated deployment pipelines and mechanisms for both physical and virtual systems, ensuring reliable and repeatable software delivery. Data-Driven Feedback & Reporting: Orchestrate processes to collect, analyze, and deliver comprehensive feedback on product stability, performance, and integration issues to segment development teams, enabling continuous improvement. Collaboration & Communication: Work closely with cross-functional teams including development, QA, security, and physical lab team to align deployment strategies, testing requirements, and environment configurations. Security & Compliance: Integrate security best practices into provisioning, deploying, and testing processes, supporting compliance with relevant standards and frameworks. This position is expected to be 100% onsite. The selected candidate will be required to work onsite in Berkeley, MO. Basic Qualifications (Required Skills/ Experien
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Fullstack Software Engineer on CLEAR’s Healthcare team, you will build and scale secure, interoperable identity and data solutions that connect patients, providers, and partners. You’ll operate at the intersection of modern web platforms, healthcare interoperability standards, and high-assurance identity systems powering frictionless, trusted healthcare experiences nationwide. A brief highlight of our tech stack: Python / React / Typescript AWS cloud What you’ll do: Design and deliver secure, scalable fullstack solutions that integrate with enterprise EHR systems and national health information exchange frameworks Build and maintain healthcare data integrations leveraging FHIR (RESTful APIs/JSON) and HL7 v2 messaging to enable compliant, real-time data exchange Develop identity resolution and patient matching capabilities using identifiers such as MRNs and NPIs to ensure integrity across disparate clinical systems Partner with Engineering, Security, Product, and Health Information Management teams to implement compliant, audit-ready workflows for regulated healthcare processes Collaborate with external vendors (e.g., Epic Technical Services) to troubleshoot integration issues, manage deployments across TST/PRD environments, and ensure production reliability How you’ll measure success: Successful delivery and stability of FHIR/HL7 integrations across healthcare partners Reduction in data integrity issues related to patient matching and identity resolution High system uptime and successful production deployments across tiered
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif
Required Skills: Core Java: Strong understanding of Java SE, including OOP concepts, data structures, and algorithms. Frameworks: Experience with popular Java frameworks such as Spring and Hibernate. Web Technologies: Knowledge of web technologies including JSP, Servlets, HTML, CSS, and JavaScript. Database Management: Proficiency in working with relational databases like MySQL, PostgreSQL, or Oracle, including SQL queries and optimization. Version Control: Experience with version control systems like Git, including branching, merging, and pull requests. Build Tools: Familiarity with build tools like Maven or Gradle. RESTful Services: Ability to design and consume RESTful web services and APIs. Testing: Experience with unit testing frameworks like JUnit or TestNG. Problem-Solving: Strong analytical and problem-solving skills with the ability to troubleshoot and resolve complex issues. Preferred Skills: Spring Boot: Experience with Spring Boot for creating microservices. ORM: Proficiency with Object-Relational Mapping (ORM) tools like Hibernate or JPA. Frontend Technologies: Basic knowledge of frontend frameworks like Angular, React, or Vue.js. Cloud Platforms: Experience with cloud platforms such as AWS, Azure, or Google Cloud. Continuous Integration/Deployment (CI/CD): Familiarity with CI/CD tools such as Jenkins or GitLab CI. Microservices Architecture: Understanding of microservices architecture and containerization with Docker. Qualifications: Education: Bachelor’s degree in Computer Science, Information Technology, or a related field. Experience: 2-4 years of professional experience in Java development. Personal Attributes: Team Player: Ability to work collaboratively in a team environment. Communication: Strong verbal and written communication skills. Attention to Detail: High attention to detail and a commitment to delivering high-quality software. Adaptability: Ability to adapt to new technologies and changing requirements.
About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability and l
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. The Cloud Efficiency team builds a unified, self-serve cloud efficiency platform along with AI skills and agents that makes spend observable, attributable, governable while driving recommendations and optimization of our cloud spend. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Design, develop, and maintain scalable platform for resource ownership registry, usage attribution, utilization measurement, and cost modeling. Build AI agents, tools and automation to enhance system monitoring, alerting, and root cause analysis. Improve and optimize data ingestion, storage, and query efficiency for cloud utilization, cost and efficiency data at scale. Collaborate with teams across Snowflake to understand attribution and observability needs and implement solutions that improve operational visibility. Contribute to open-source and industry best practices in monitoring and distributed systems monitoring. Ensure high availability, reliability, and performance of team-managed platforms by participating in on-call rotations and incident management. Partner with Finance, Product and Engineering
Location Details: Canada, Remote At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) , and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability a
Location: Mumbai, India Experience: 8–10 Years Qualification: CA / MBA (Finance) / CFA or equivalent About the Role- We are looking for an experienced Treasury professional to join our Corporate Treasury & Capital Markets team. The role will be responsible for managing the organization's funding, liquidity, banking relationships, capital market borrowings, treasury operations, and regulatory compliance while ensuring efficient cash management and optimal utilization of financial resources. The ideal candidate will have strong experience in debt capital markets, corporate treasury operations, banking relationships, SAP, and regulatory compliance. Key Responsibilities 1. Capital Markets & Debt Management Structure, execute, and manage short-term and long-term debt instruments including: Commercial Papers (CPs) Non-Convertible Debentures (NCDs) Term Loans Coordinate with credit rating agencies, merchant bankers, debenture trustees, institutional investors, and lending banks. Support fund-raising initiatives and debt optimization strategies. Manage security creation, charge filings (ROC/CERSAI), and all related financing documentation. Monitor debt portfolio, repayment schedules, and borrowing costs. 2. Treasury Investments & Liquidity Management Manage daily liquidity and cash positioning across all bank accounts. Optimize deployment of surplus funds through: Money Market Instruments Mutual Funds Fixed Income Investments Prepare and monitor daily cash flow and liquidity reports. Develop and enhance working capital and cash flow forecasting models. Ensure efficient allocation of funds while minimizing idle cash balances. 3. Banking Relationships & Treasury Operations Manage strategic relationships with banks and financial institutions. Negotiate banking facilities, pricing, and credit limits. Oversee Corporate Banking platforms including: Host-to-Host (H2H) Cash Management Services (CMS) Online Banking Portals Administer banking access controls, user rights, and s
Get new platform deployment management lead jobs by email
Daily job updates · Unsubscribe anytime