Jobs in United States

Platform Deployment Management Lead in United States

3,022 active opportunities · Updated October 2026

Explore current platform deployment management lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

L
📍 Hampton, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

The Defense Sector at Leidos is seeking a motivated TS/SCI cleared Network Administrator to support the installation, configuration, and day-to-day management of enterprise network infrastructure. This role is an excellent opportunity for an early-career network professional to gain hands-on experience with routing and switching platforms, including Session Smart Router (SSR) / 128 Technology SD-WAN solutions, in a structured and security-conscious environment. The ideal candidate demonstrates a solid foundation in networking fundamentals, a willingness to learn vendor-specific technologies, and the discipline to operate within DoD network standards. The job duties will be performed daily on site at Langley Air Force Base, VA. ​ Roles and Responsibilities: ​Assist in the configuration, deployment, and ongoing management of routers, switches, and Session Smart Router (SSR) appliances across enterprise and edge network environments. ​Support the design and implementation of routing policies, service policies, and traffic steering configurations on SSR/128 Technology platforms under senior engineer guidance. ​Perform LAN switching administration — including VLAN configuration, spanning tree, trunking, and port security — on Juniper EX Series and/or Cisco Catalyst platforms. ​Assist with the configuration and troubleshooting of routing protocols including OSPF and BGP (eBGP and iBGP) across enterprise WAN and data center environments. ​Monitor network health, availability, latency, and throughput using network management tools; escalate anomalies and assist in root cause analysis. ​Support configuration and maintenance of firewall rules, access control lists (ACLs), IPsec VPN tunnels, and other network security controls. ​Execute software and firmware upgrades, patch management, and lifecycle maintenance activiti

PythonAnsible
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you will be doing: Design and develop firmware solutions for manageability and observability of data center servers. Actively participate in hardware bring-up activities, OOB firmware development, protocol stacks (Redfish, PLDM, MCTP, NSM) and hardware-software co-design for Cloud Service Provider deployments. Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Deep expertise in embedded firmware, server management controllers, and hardware bring-up with proven track record of shipping production BMC solutions Strong knowledge of DMTF protocols (Redfish, IPMI, PLDM, MCTP, SPDM), telemetry frameworks, and out-of-band management architectures Expert-level skills in C/C&

Artificial IntelligenceAI
C
📍 Woonsocket, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. The Release Engineer is responsible for implementing and supporting automated software delivery processes that enable reliable, secure, and repeatable deployments across development, test, non-production, and production environments. This role serves as a technical bridge between Software Engineering, Quality Engineering, Infrastructure, and Operations teams to improve deployment speed, stability, and operational efficiency. Key Responsibilities Develop, support, and maintain automated CI/CD pipelines to streamline application delivery. Standardize build, release, and deployment processes across applications and platforms. Establish repeatable, auditable, and traceable release management practices to ensure deployment consistency and compliance. Manage and support Git-based source control, branching strategies, and release workflows. Integrate automated testing, quality gates, security controls, and compliance checks into deploym

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal

AWSRestAIRust
B
📍 Berkeley, United States
✓ Quality checkedCompany trend +515.8%

Associate and Mid-Level Software Engineers Company: The Boeing Company The Boeing Company is looking for an Associate or Mid-Level Software Engineer to join our team in Berkeley, MO. Position Responsibilities: Lab Environment Provisioning: Design, provision, and maintain scalable lab environments on-premises and on cloud platforms (Azure, AWS) using Infrastructure as Code (IaC) tools such as Terraform and Ansible. Automated Integration Testing Orchestration: Develop and manage automated integration testing workflows that aggregate inputs from multiple program segments, ensuring comprehensive test coverage and timely feedback. Deployment Management: Manage and optimize automated deployment pipelines and mechanisms for both physical and virtual systems, ensuring reliable and repeatable software delivery. Data-Driven Feedback & Reporting: Orchestrate processes to collect, analyze, and deliver comprehensive feedback on product stability, performance, and integration issues to segment development teams, enabling continuous improvement. Collaboration & Communication: Work closely with cross-functional teams including development, QA, security, and physical lab team to align deployment strategies, testing requirements, and environment configurations. Security & Compliance: Integrate security best practices into provisioning, deploying, and testing processes, supporting compliance with relevant standards and frameworks. This position is expected to be 100% onsite. The selected candidate will be required to work onsite in Berkeley, MO. Basic Qualifications (Required Skills/ Experien

PythonAWSAzureDocker
C-
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Fullstack Software Engineer on CLEAR’s Healthcare team, you will build and scale secure, interoperable identity and data solutions that connect patients, providers, and partners. You’ll operate at the intersection of modern web platforms, healthcare interoperability standards, and high-assurance identity systems powering frictionless, trusted healthcare experiences nationwide. A brief highlight of our tech stack: Python / React / Typescript AWS cloud What you’ll do: Design and deliver secure, scalable fullstack solutions that integrate with enterprise EHR systems and national health information exchange frameworks Build and maintain healthcare data integrations leveraging FHIR (RESTful APIs/JSON) and HL7 v2 messaging to enable compliant, real-time data exchange Develop identity resolution and patient matching capabilities using identifiers such as MRNs and NPIs to ensure integrity across disparate clinical systems Partner with Engineering, Security, Product, and Health Information Management teams to implement compliant, audit-ready workflows for regulated healthcare processes Collaborate with external vendors (e.g., Epic Technical Services) to troubleshoot integration issues, manage deployments across TST/PRD environments, and ensure production reliability How you’ll measure success: Successful delivery and stability of FHIR/HL7 integrations across healthcare partners Reduction in data integrity issues related to patient matching and identity resolution High system uptime and successful production deployments across tiered

TypeScriptPythonReactAWS
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

AWSKubernetesLinuxRest
S
📍 Bellevue, Washington, United States· Full-time
✓ High-confidence listingCompany trend -92.9%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. The Cloud Efficiency team builds a unified, self-serve cloud efficiency platform along with AI skills and agents that makes spend observable, attributable, governable while driving recommendations and optimization of our cloud spend. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Design, develop, and maintain scalable platform for resource ownership registry, usage attribution, utilization measurement, and cost modeling. Build AI agents, tools and automation to enhance system monitoring, alerting, and root cause analysis. Improve and optimize data ingestion, storage, and query efficiency for cloud utilization, cost and efficiency data at scale. Collaborate with teams across Snowflake to understand attribution and observability needs and implement solutions that improve operational visibility. Contribute to open-source and industry best practices in monitoring and distributed systems monitoring. Ensure high availability, reliability, and performance of team-managed platforms by participating in on-call rotations and incident management. Partner with Finance, Product and Engineering

PythonJavaAWSAzure
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%

From $230K/yr

Quick readStrong listing-quality and freshness signals

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

PythonAWSKubernetesGit
A
📍 San Francisco, CA, United States
✓ Quality checkedCompany trend -100%

The Opportunity Typography is central to how ideas are communicated. If you're passionate about beautifully created design, have deep curiosity about what AI can do, and take personal responsibility for creating products that people love; then this may be the role for you. Adobe Fonts supports millions of creatives in choosing and using typefaces across fonts.adobe.com, Express, Photoshop, Illustrator, Acrobat, and more Creative Cloud platforms. Our Internal Services team provides the platform engineering and deployment backbone for all of these. We manage CI/CD, deployment approaches, the services and data layers our engineers depend on, our observability and security stance, and increasingly the agentic tools that transform how our entire organization delivers software. We're seeking a Senior Software Development Engineer to lead this exciting journey in our San Francisco location. What you'll do Own and evolve our deployment platform. Lead strategy for CI/CD, PR environments, and release safety across a mixed fleet that includes containerized services, serverless services, and static front ends. Build the foundation for AI-accelerated development. Help build our agent factory and grow our internal agentic toolkit and skill library. Ship inference applications at scale. Take greenfield services from spec to production and standardize our ML/inference footprint. Modernize our services for the AI era. Identify where an existing service is held back by its current build and lead the fix. Rethink our security posture for agentic threats. Lead how we secure autonomous agents and their tool use. Expose Adobe Fonts to the agentic ecosystem. Extend our Model Context Protocol (MCP) surface and conversational, intent-based font discovery. Work higher up the stack, too. Contribute directly to search, browse, discovery, and the customer-facing experie

JavaScriptTypeScriptPythonReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Growth team drives user and revenue growth across ChatGPT’s consumer and business segments as well as other OpenAI products worldwide. We operate across the full funnel - from awareness and acquisition through activation, retention, and expansion - using a combination of global performance marketing, AI-powered workflows, in-product optimization, insights, experimentation, and creative ops engineering. About the Role We are hiring a Lifecycle Lead to build the company-wide owned-channel capability that helps teams reach users with relevant, timely, and trustworthy experiences. This senior, hands-on leader will set the lifecycle strategy, partner with Engineering to build the orchestration and deployment platform, and establish the operating model that allows teams across the company to launch and improve evergreen programs safely at scale. You will sit at the intersection of platform, product, and campaign strategy. You will partner with Engineering, Product, Data Science, and Analytics on the underlying systems, and with Product Marketing Managers and other client teams to design journeys that help new, active, and returning users reach value and build durable habits. In this role, you will: Partner with product to set the company-wide vision, roadmap, and operating model for lifecycle and owned-channel engagement. Partner with Engineering, Product, Data Science, and Analytics to shape the tooling and infrastructure for identity, audiences, eligibility, consent, triggers, orchestration, decisioning, frequency, experimentation, localization, quality assurance, and observability. Define scalable deployment workflows—including self-service and centrally supported paths, intake, templates, approvals, governance, service levels, and incident response—so teams across the company can launch safely and efficiently. Partner with Product Marketing Managers and other client teams to translate audience, product, and business goals into evergreen journey stra

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As an early member of Baseten's Platform Team, you will be pivotal in building internal infrastructure to support our engineering organization. You will own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across Baseten to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. If you are passionate about elegant solutions—like streamlined monorepos, lightning-fast CI pipelines, and thoughtfully designed shared libraries—you'll thrive at Baseten. RESPONSIBILITIES Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces,

PythonKubernetesGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Infrastructure team at OpenAI is responsible for building systems that enable training and deployment of highly useful AI agents, both internally and for the world. We work hand-in-hand with researchers to design and scale the environment in which agentic models are trained – providing a workspace for AI models to execute code, debug issues, and develop software just as human SWEs do. Our training environment for agentic models operates at an extremely high scale and has the flexibility to emulate any environment in which an agent might work. At the same time, our team builds and maintains OpenAI’s core platform for the deployment and execution of agents in production. Our systems power products such as Codex, Operator, tool use in ChatGPT, and future agentic products. Some of the most challenging technical problems in scaling the capabilities and utility of agents and agentic models lie in the infrastructure layer – and our team is focused on building the research and production systems that enable OpenAI to train the most capable models in the world, and maximize the utility of our agentic products for users around the world. About the Role As a Software Engineer on the Agent Infrastructure team, you will have the opportunity to work closely with both research and product at OpenAI - building and scaling systems to train highly capable agentic models, and building the platform and integrations to launch new agents to hundreds of millions of users worldwide. Your work will consist of both building new capabilities - standing up the infrastructure and integrations needed to train more complex agentic models - and rapidly scaling these new capabilities to some of the largest compute clusters in the world. At the same time, you’ll be instrumental to the launch of agentic products at OpenAI - building, maintaining, and scaling the production platform on which all agents run. We’re looking for people with deep experience building AI infrastructu

AWSKubernetesRestMachine Learning
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. Join NVIDIA's NIM team and be part of an exceptionally ambitious project in Santa Clara, CA! As a Senior Software Engineer, NIM Tools, you will have the remarkable opportunity to build a groundbreaking model customization and deployment lifecycle platform from inception. This isn't just another feature team—you will be defining the structure for a new product surface accessed by ISVs and CSPs internationally. Your work will empower customers to take models from selection through fine-tuning, evaluation, deployment, and compliance flawlessly. What you'll be doing: Compose and build the fine-tuning handoff pipeline, including LoRA adapter repackaging, re-quantization, and re-validation into NIM. Develop the evaluation harness, ensuring models meet our high standards. Implement the observability and attestation layer to produce auditable compliance artifacts. Work in close partnership with ISVs and CSPs to roll out NVIDIA NIMs on a large scale. Define and improve durable platform APIs, steering clear of one-off integrations. Ensure flawless completion of projects through strict attention to detail and proven methodologies. Wha

🔔

Get new platform deployment management lead jobs in United States by email

Daily job updates · Unsubscribe anytime