About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a System Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate low-level system software for the manageability and health of OpenAI's first-party AI hardware systems. You will work across BMC, Linux, firmware interfaces, automation infra, boot and recovery, hardware diagnostics, telemetry, host and platform drivers, network software interfaces, and manufacturing and fleet readiness. A major part of this role is owning the acceptance path for partner-delivered system software: defining requirements, reviewing code and artifacts, reproducing builds, building tests, pushing fixes, and producing the evidence needed for launch decisions. This role is hands-on and high-ownership. You will write and review low-level software, debug issues across hardware and software boundaries, build infra and automation to test and manage devices in lab, guide partner deliverables, build validation evidence, and help carry platforms from bring-up through production deployment. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Design, develop, and maintain low-level firmware and system software for first-party AI hardware manageability, including BMC software, Redfish services, gNMI telemetry, firmware update and recovery flows, BIOS/UEFI interactions, platform drivers, and hardware diagnostics. Own integration and acceptance of partner and ve
Jobs in United States
Deployment Strategist in United States
636 active opportunities · Updated October 2026
Showing
15 jobs
Explore current deployment strategist jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
About the Team Safety Systems works to ensure OpenAI’s most capable models can be developed and deployed responsibly. Our work spans evaluations, safeguards, red teaming, deployment decisions, and the systems that help OpenAI understand and reduce risk as models become more capable and widely used. Within Safety Systems, the Trustworthy AI team is growing its safety transparency function: a practice focused on helping external audiences understand OpenAI’s technical safety work with greater clarity, rigor, and continuity. We create and improve the public artifacts that explain how our systems are evaluated for safety, what safeguards we build, what decisions we make, and where uncertainty remains. This work includes system cards, the Deployment Safety Hub, safety-related blogs, public governance documents, and other outputs that communicate technical safety topics to external audiences. It also includes building new ways to make technical safety information easier to understand, navigate, and use—including AI-assisted workflows, data visualizations, and interactive tools that make complex technical work more legible over time. About the Role We are looking for a Safety Transparency Editor to own the editorial quality of key safety transparency artifacts and systems. This is a hands-on role for someone who can write crystal-clear, pitch-perfect explanations of the hardest and highest-stakes technical safety topics that OpenAI tackles, and who can lean into AI to build systems that help the broader organization do this work better. Your core responsibility is to shape and execute how our technical safety work is externally communicated: identifying the narrative thread, exercising judgment about which details matter, determining where additional context, explanation, or supporting evidence is needed, translating complexity without sacrificing precision, and helping external audiences understand both the safety measures we’ve taken and the uncertainties that remain. To
About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. We believe that achieving our goal requires real world deployment and iteratively updating based on what we learn. The Protection Scientist Engineer, Integrity team supports this by identifying and investigating misuses of our products – especially new types of abuse. This enables our partner teams to develop data-backed product policies and build scaled safety mitigations. Precisely understanding abuse allows us to safely enable users to build useful things with our products. About the Role Protection Science Engineering is an interdisciplinary role mixing data science, machine learning, investigation, and policy/protocol development. As a Protection Scientist Engineer within Integrity and Investigations, you will be responsible for designing and building systems to proactively identify and enforce on abuse on OpenAI’s products. This includes ensuring we have robust abuse monitoring in place for new products, sustaining monitoring for existing products, and prototyping and incubating systems of defense against our highest risk harms. You will also respond to and investigate critical escalations, especially those that are not caught by our existing safety systems. This will require expert understanding of our products and data, and involves working cross-functionally with product, policy, and engineering teams. This role can be based in either our San Francisco, or NY office and includes participation in an on-call rotation that will involve resolving urgent escalations outside of normal work hours. Some investigations may involve sensitive content, including sexual, violent, or otherwise-disturbing material. In this role, you will: Scope and implement abuse monitoring requirements for new product launches. Improve processes to sustain monitoring operations for existing products, including developing approaches to automate monitoring subtasks. Prototyp
About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Learn more about OpenAI’s approach to safety. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. About the Role At OpenAI, we're dedicated to advancing artificial intelligence, and we know that creating a secure and reliable platform is vital to our mission. That's why we're seeking a software engineer to help us build out our trust and safety capabilities. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. Your Responsibilities: Architect, build, and maintain anti-abuse and content moderation infrastructure designed to protect us and end users from unwanted behavior. Work closely with our other engineers and researchers to utilize both industry standard and novel AI techniques to measure, monitor and improve AI models’ alignment to human values. . Diagnose and remediate active incidents on the platform and build new tooling and infrastructure that address the root causes of system failure. You might thrive in this role if: You have built and run production services in a high growth, rapidly scaling environment. You can debug live issues and restore systems quickly. You have worked on content safety, fraud, or abuse, or are motivated and excited to work on present-day (“now-term”) AI safety. You have experience with Python or with modern languages such as C++, Rust, or Go, and are able to quickly ramp up on Py
About the Team OpenAI’s AI Success Engineer team partners with the world’s most ambitious government & partner organizations to translate cutting edge AI into real business and mission impact for governments of all levels from Local, State, Federal, and International. We guide customers and users journey from the first time they try ChatGPT Enterprise, automate a workflow, develop and execute a new skill, and create their first agent to scaled enterprise adoption of ChatGPT, Codex, our API and other novel capabilities. Our work spans technical integration and enablement, workflow transformation, inspiring and upskilling AI literacy and confidence across the workforce, sustained program, product and new capability delivery. Most importantly, we help each member of our customer's workforce, their teams, programs and missions meet their total potential. Our government customers have vital missions, and we must meet them with game-changing technology. Every engagement is an opportunity to shape how AI changes work, productivity, and innovation. This role sits at the center of that mission. About the Role Governments work at a scale that is truly exponential on missions that are of critical importance to people, communities and nations. The AI Success Engineer role is the primary post-sales relationship for OpenAI’s most important customers. You are responsible for the end-to-end account management of critical Government and Partner customers. You will be helping Government Leaders/Partners appropriately and effectively use AI for their mission, while simultaneously investing in ensuring their people are AI-enabled and ready to advance positive outcomes that their constituents depend on them for. You will drive: the impact of our tools on their mission, account health and adoption, ensuring technical readiness, creating and executing on the deployment strategy, enabling, educating and training their workforce, identifying new use cases and upsell opportunities, and d
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Job Description Serves as a Technical Digital Product Implementation Manager for Aetna Health Customizations supporting plan sponsors and Book of Business products. Candidate will lead and oversee the deployment and integration of advanced digital technologies, including SSO, API, and content management systems. Role requires autonomously driving successful implementation while managing multiple deliveries and collaborating with several business partners. This position requires strong communication, project management, and technical skills combined with healthcare industry knowledge to implement digital enhancements that improve the member experience, create operational efficiency, and ensure compliance with industry regulations. Key Responsibilities: Independently lead the end-to-end implementation of digital technology solutions, including SSO and API integrations, in conjunction with internal and external partners. Own all portions of project delivery, including; requirement gathering, project documentation, facilitation and collaboration efforts with all impacted parties, testing, release, and support. Collaborate with business stakeholders to understand their needs, define project scope, and ensure alignment with organizational goals. Communicate project progress, risks, and/or blockers, to progress project and show transparency in approac
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Job Description The Digital Product Implementation Manager supports Aetna Health customizations and leads the definition, deployment and integration of advanced digital capabilities across Aetna’s customers. This role independently drives successful implementations while managing multiple deliveries, collaborating with a range of business partners, and working directly with customers throughout the implementation process. You will sit at the intersection of the Technology team, the Business, and the Customer to ensure we are building best-in-class digital experiences for our members and clients. As an enthusiastic, self-led professional, you will leverage your passion for doing what’s right for our members to deliver the right solutions to meet business and market needs. You have superior program management and communication skills, strong attention to detail and are able to manage multiple priorities at once. You thrive on collaboration and working across teams. You are a self-starter who can expertly navigate complex organizations. You have a digital-first, data-driven mindset. You are quick to adapt to business and customer needs to deliver digital enhancements that improve the member experience, create operational efficiency, and ensure compliance with industry regulations. Key Responsibilities Independently lead the end-to-end customer implementation
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Rack Power Engineer with deep expertise in high-power conversion and distribution to design, qualify, and support power systems for AI supercomputers. You will own rack power solutions—including power shelves, AC/DC rectifiers, power supply units (PSUs), power management controllers (PMCs), and high-current distribution—from requirements and supplier development through deployment. You will also monitor fleet rack power health, lead debugging and root-cause investigations, and drive improvements into hardware, firmware, and qualification coverage. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own rack power architecture and requirements for high-power AI supercomputing systems, including power budgets, AC input interfaces, DC distribution, redundancy, efficiency, serviceability, and integration with data center infrastructure. Drive the design and supplier development of power shelves, rectifiers, PSUs, PMCs, busbars, connectors, and protection circuits. Review electrical designs and control behavior, and evaluate performance, cost, reliability, and availability trade-offs. Define and execute component, shelf, and rack qualification plans covering load transients, current sharing, hot-swap, startup and shutdown, redundancy failover, fault protection and recovery, thermal limits, and AC disturbances and ride-through
Senior Ground-based Midcourse Defense (GMD) Maintenance Operations Center (GMOC) Controller Company: Boeing Aerospace Operations Boeing Global Services (BGS) is seeking a Senior Ground-based Midcourse Defense (GMD) Maintenance Operations Center (GMOC) Controller located at the Missile Defense Integration and Operation Center (MDIOC) at Schriever Space Force Base, Colorado. Must be willing to work variable shifts, including weekends and overtime. Some positions may also be rotating shifts. You are applying for a Field Operations position where our employees must be willing and able (potentially short notice) to travel to a variety of locations (domestic and/or international) to meet our customers’ unique requirements. Field Operations employees may be required to relocate to another Field Operations location based on customer and/or management requirements. Specific contracts supported by Field Operations could require employees to perform a deployment to a US DOD contingency operation. Employees associated with these specific contracts must be able to pass military contractor health standards for deployment. Position Responsibilities: The position frequently uses and applies operational concept analysis standards, principles, theories, concepts, and techniques Crew Lead that research, analyze, and compile complex technical data for the GMD program in the operational and test environments to optimize effectiveness over the program lifecycle Maintain configuration control of the GMD System 24/7/365 days a year supporting the Missile Defense Agency and warfighter community Perform asset management functio
As one of the technology industry's most desirable employers, NVIDIA has been redefining accelerated computing, computer graphics and leading the Artificial Intelligence revolution. NVIDIA's innovation is fueled by its great technology—and amazing people. We are seeking a Senior Silicon and System Product Lead to influence, innovate and take our next generation products to the market. As part of the Silicon Solutions Team, we architect and deliver groundbreaking system solutions that integrate all aspects of the system from silicon design, software design to operations and final deployment in multiple market segments that NVIDIA serves. This position offers an unique opportunity to collaborate with multiple organizations in the company and grow your career in a high impact role. We need a passionate, hard-working and creative individual to lead the products all the way from market analysis to delivering the features on the final product. What you'll be doing: Drive product performance and power targets, trade-off features/configurations and provide innovative solutions to complex silicon and system level problems. Evaluate new market segments and use cases; translate market requirements to engineering problem statements and metrics. Innovate Performance, power, yield and quality optimizations and features for the world’s fastest power-shipping products in the GPU and SoC market segments spanning gaming, automotive, datacenter and DL/AI. Develop methodologies and requirements for multi-functional teams to drive silicon and system product features to production. Incorporate productization feedback to improve the next generation. Lead the team for feature requirements and schedule from architecture to silicon phase of projects. Work alongside system architects, designers, marketing teams, chip and board designers, software/firmware engineers, HW/S
Are you ready to contribute to world-class innovation and push the boundaries of what's possible? At NVIDIA, you'll have the opportunity to be part of a team that is driving groundbreaking impacts across various markets. As a Thermal Solutions Development Engineer, you will play a pivotal role in our Silicon Codesign Group, transforming thermal solution concepts into lab-ready builds and beyond. What you will be doing: Build thermal solutions for engineering characterization and validation of next-gen GPU/SOC products, ensuring flawless delivery from concept to lab. Drive end-to-end development and deployment of thermal solutions, collaborating with internal teams and external vendors on build requirements, prototype evaluation, test system integration, and software automation. Improve thermal design processes by incorporating feedback and findings, developing workflow and maintaining our world-class standards. Work closely with system architects, chip and board designers, and software/firmware engineers in a dynamic and high-energy environment to bring industry-defining products to market. Apply AI-enabled approaches and AI tools to accelerate design iteration, test planning, and characterization/validation triage (e.g., requirements/spec summarization, experiment prioritization, log/telemetry summarization, anomaly/outlier detection), improving cycle time, coverage, and traceability while validating outputs against physics, specs, and lab measurements. Partner with AI/tooling teams as the thermal domain SME to define use-cases, success criteria, and evaluation methods; provide feedback to improve tool reliability and usability. What we need to see:
$180K – $190K/yr
Sapience AI is the collective intelligence platform for professional communities. We sit above the CRMs, AMS platforms, and knowledge bases that organizations already run, and we turn the expertise scattered across them into something every member can search, act on, and share. The intelligence a community needs is already inside it. Most organizations just cannot reach it. Knowledge lives in silos, in legacy systems, in the heads of a few experts, and in fragmented records no one can connect. We change that. Our work is grounded in four commitments: technology elevates people and never replaces them, the best expertise is already inside the community, everything is built on trust, and every deployment is purpose-driven for the organization it serves. Let’s achieve more, together. Where this role sits This role owns a product area end to end. Where the product marketer is the market lens, you are the person accountable for what we build, how it works, and whether it ships and succeeds. You lead a cross-functional team of engineers, designers, and applied AI specialists to turn strategy into a product members rely on. You own the roadmap for your area, the decisions inside it, and the outcomes it produces. You work at the center of the product organization, translating a clear point of view into shipped work that moves MINERVA and the COGENT architecture forward through each GR/GA milestone. Why this role exists In an AI market where capabilities can be copied in weeks, the advantage goes to teams that decide well and ship fast. Someone has to hold the point of view, make the calls, and keep a team building the right thing. Collective intelligence is hard product work: reasoning over messy knowledge, earning member trust, and fitting into how communities already operate. It needs an owner who can hold both the ambition and the details. The Senior Product Manager is that owner. You set direction for your area, make the hard prioritization calls, and are accountable fo
The Defense Sector at Leidos is seeking a motivated TS/SCI cleared Network Administrator to support the installation, configuration, and day-to-day management of enterprise network infrastructure. This role is an excellent opportunity for an early-career network professional to gain hands-on experience with routing and switching platforms, including Session Smart Router (SSR) / 128 Technology SD-WAN solutions, in a structured and security-conscious environment. The ideal candidate demonstrates a solid foundation in networking fundamentals, a willingness to learn vendor-specific technologies, and the discipline to operate within DoD network standards. The job duties will be performed daily on site at Langley Air Force Base, VA. Roles and Responsibilities: Assist in the configuration, deployment, and ongoing management of routers, switches, and Session Smart Router (SSR) appliances across enterprise and edge network environments. Support the design and implementation of routing policies, service policies, and traffic steering configurations on SSR/128 Technology platforms under senior engineer guidance. Perform LAN switching administration — including VLAN configuration, spanning tree, trunking, and port security — on Juniper EX Series and/or Cisco Catalyst platforms. Assist with the configuration and troubleshooting of routing protocols including OSPF and BGP (eBGP and iBGP) across enterprise WAN and data center environments. Monitor network health, availability, latency, and throughput using network management tools; escalate anomalies and assist in root cause analysis. Support configuration and maintenance of firewall rules, access control lists (ACLs), IPsec VPN tunnels, and other network security controls. Execute software and firmware upgrades, patch management, and lifecycle maintenance activiti
Other cities to consider
More places hiring for this role
Get new deployment strategist jobs in United States by email
Daily job updates · Unsubscribe anytime