The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur
Jobiba hiring network
Deployment Strategist Lead Jobs
1,782 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current deployment strategist lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Safely delivering increasingly capable AI systems requires scalable technical safeguards, clear ownership of emerging risks, rigorous deployment readiness, and close coordination across research, engineering, product, operations, legal, policy, and external partners. Our Technical Program Managers lead complex, high-stakes initiatives that turn safety commitments into deployed systems and measurable outcomes. We work across model development, infrastructure, product, and operational response to help ensure our technology is deployed responsibly and cannot be used to cause serious real-world harm. About the Role We’re seeking Technical Program Managers to drive complex product, platform, and safety initiatives across ChatGPT, API, enterprise, and related deployment environments. These roles operate at the intersection of technical strategy and execution: you will turn safety and product priorities into actionable plans, influence architectural and operational decisions, and deliver durable capabilities across model, infrastructure, application, and platform layers. Depending on the role, you may enable sensitive or high-impact model deployments, integrate safeguards into cloud and API platforms, prevent violent misuse and other serious harms, improve detection and enforcement systems, create platform solutions for safety or establish new programs as risks evolve. You will partner deeply with engineers, researchers, product managers, and operational teams while communicating technical tradeoffs and program decisions to senior leadership. You bring technical fluency, product judgment, and a strong execution record. You’re comfortable navigating ambiguity, advocating for users and developers, balancing safety with model usefulness, and leading cross-functional work with urgency, rigor, and empathy. Specific focus areas and scope will vary by opening and level. In
About the team OpenAI’s Forward Deployed Engineering team partners with healthcare organizations to deploy production AI systems across clinical, operational, and member-facing workflows. We work at the boundary of customer deployment and core platform development, using customer engagements to define repeatable architectures, evaluations, integrations, and operating standards for complex, regulated healthcare environments. About the role We are hiring a Forward Deployed Engineer (FDE) to own end-to-end deployments of our models within healthcare organizations, including payers, providers, health systems, and healthcare technology companies. You will lead technical discovery, architecture, implementation, evaluation, productionization, and handoff, translating complex customer workflows, data, infrastructure, and regulatory constraints into production AI systems. You will measure success through production adoption, measurable workflow impact, and evaluation loops that establish customer-specific benchmarks, acceptance criteria, and launch readiness. You’ll collaborate directly with customer technical and operational teams, alongside internal Business, Research, Product, Engineering, and Security partners, to deliver solutions and translate deployment learnings into product improvements. This role owns the technical solution; ownership of the commercial or executive relationship is not required. This role is based New York City. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% is required. In this role you will Own the technical solution end to end, from customer discovery and workflow scoping through architecture, hands-on implementation, evaluation, production deployment, adoption, and handoff. Partner credibly with customer engineers, operators, and domain experts to frame ambiguous problems, define scope, and translate payer, provider, or health-system workflows into technical requirements and measurab
About the team OpenAI’s Forward Deployed Engineering team partners with healthcare organizations to deploy production AI systems across clinical, operational, and member-facing workflows. We work at the boundary of customer deployment and core platform development, using customer engagements to define repeatable architectures, evaluations, integrations, and operating standards for complex, regulated healthcare environments. About the role We are hiring a Forward Deployed Engineer (FDE) to own end-to-end deployments of our models within healthcare organizations, including payers, providers, health systems, and healthcare technology companies. You will lead technical discovery, architecture, implementation, evaluation, productionization, and handoff, translating complex customer workflows, data, infrastructure, and regulatory constraints into production AI systems. You will measure success through production adoption, measurable workflow impact, and evaluation loops that establish customer-specific benchmarks, acceptance criteria, and launch readiness. You’ll collaborate directly with customer technical and operational teams, alongside internal Business, Research, Product, Engineering, and Security partners, to deliver solutions and translate deployment learnings into product improvements. This role owns the technical solution; ownership of the commercial or executive relationship is not required. This role is based in Seattle. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% is required. In this role you will Own the technical solution end to end, from customer discovery and workflow scoping through architecture, hands-on implementation, evaluation, production deployment, adoption, and handoff. Partner credibly with customer engineers, operators, and domain experts to frame ambiguous problems, define scope, and translate payer, provider, or health-system workflows into technical requirements and measurable
About the Team OpenAI’s Forward Deployed Engineering team partners with healthcare organizations to deploy production AI systems across clinical, operational, and member-facing workflows. We work at the boundary of customer deployment and core platform development, using customer engagements to define repeatable architectures, evaluations, integrations, and operating standards for complex, regulated healthcare environments. About the Role We are hiring a Forward Deployed Engineer (FDE) to own end-to-end deployments of our models within healthcare organizations, including payers, providers, health systems, and healthcare technology companies. You will lead technical discovery, architecture, implementation, evaluation, productionization, and handoff, translating complex customer workflows, data, infrastructure, and regulatory constraints into production AI systems. You will measure success through production adoption, measurable workflow impact, and evaluation loops that establish customer-specific benchmarks, acceptance criteria, and launch readiness. You’ll collaborate directly with customer technical and operational teams, alongside internal Business, Research, Product, Engineering, and Security partners, to deliver solutions and translate deployment learnings into product improvements. This role owns the technical solution; ownership of the commercial or executive relationship is not required. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% is required. In this role, you will: Own the technical solution end to end, from customer discovery and workflow scoping through architecture, hands-on implementation, evaluation, production deployment, adoption, and handoff. Partner credibly with customer engineers, operators, and domain experts to frame ambiguous problems, define scope, and translate payer, provider, or health-system workflows into technical requirements and mea
About the Team The ChatGPT Model Flywheel team unified goal is to transform model advancements into great ChatGPT user experiences through reliable serving, rapid experimentation, safe deployment, and continuous improvement. Team Focus Areas Model Experimentation: Enable rapid, safe model validation for ChatGPT and Codex products through experiment automation and lifecycle management. Model Deployment: Ensure safe, scalable deployment of model capabilities with robust rollout and operational tooling. Automate capacity management and incorporate platform-wide health monitors. Model Measurement: Build comprehensive evaluation and measurement systems for model quality, from user signals to launch scorecards. Improve end-to-end feedback loops for continual model improvement. Key Partnerships Collaborate cross-functionally with teams including Model Measurement DS, Research, Codex, Fleet, Inference, and API. In this role, you will: Elevate and consolidate ChatGPT’s harness, context management, and system prompt frameworks. Drive expansion and improvement of multi-tier model experiences. Support and scale self-serve experiment capabilities and automated guardrails. Lead model rollout automation, capacity management, and health monitoring. Shape end-to-end measurement systems (evals, grader signals, user feedback, etc.). You might thrive in this role if you have: Proven experience leading engineering teams in complex, cross-functional environments. Demonstrated success shipping production systems at scale (ideally for AI or large backend services). Deep understanding of model-driven product development, deployment lifecycle, and measurement tooling. Excellent communication and collaboration skills—experience interfacing directly with engineering, research, and product stakeholders. Prior involvement with large language models, distributed infrastructure, or experimentation platforms is a plus. Why Work With Us Tackle highly impactful technical challenges at the cutting edg
About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We are looking for an experienced Research Engineer to work on retrieval & search problems across our API and ChatGPT. As the AI landscape has evolved over the last few years, retrieval & search have emerged as key use cases for our models, and we are investing in ensuring that we can offer these search-based product experiences for our users. You will be at the center of our retrieval & search efforts as a company, and the progress you drive here will reach millions of end users. In this role, you will: Work on retrieval & search algorithms and methodologies in close collaboration with our research team, including problems in such domains as document search, enterprise search, knowledge retrieval, and web-scale search. Deploy these search methodologies into production in both the API and ChatGPT to be used by millions of end users. Explore novel research topics in retrieval & search that may inform our product strategy in the medium and long term. Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world You might thrive in this role if you: Have extensive prior experience building and maintaining production machine learning systems. Have prior experience working with vector databases, search indices, or other data stores for search and retrieval use cases Have prior experience building and iterating on internet-scale search systems Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done Have the ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or de
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We embed deeply with users to solve high-leverage problems. We move quickly from prototype to deployment and surface patterns that shape the platform. We operate at the intersection of customer delivery and core development. We work closely with Product, Research, and Go-To-Market (GTM). About the role Forward Deployed Engineers lead complex deployments of frontier models in production. You will embed with customers where model performance matters, delivery is urgent, and ambiguity is the default. You will use this to map their problems,You will use this to map their problems, structure delivery, and ship fast. You will scope, sequence, and build full-stack solutions that create measurable value. You will also drive clarity across internal and external teams. You will identify reusable patterns and share field signal that influences the roadmap. Success in this role means owning the delivery state across workstreams. You will hold the bar on quality and pace and help OpenAI learn through execution. This role is based in Abu Dhabi. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% is required. In this role you will Own technical delivery across multiple deployments from first prototype to stable production Build full-stack systems that deliver customer value and sharpen how we learn Embed closely with customer teams, understand their needs, and guide adoption of what you build Scope work, sequence delivery, and remove blockers early Make trade-offs between scope, speed, and quality; adjust plans to protect delivery Contribute directly in the code when progress or clarity depends on it Codify working patterns into tools, playbooks, or building blocks that others can use Share field feedback that helps Research and Product understand where the models succeed and where they c
About the team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the role: We're seeking a Data Engineer to take the lead in building our data pipelines and core tables for OpenAI. These pipelines are crucial for powering analyses, safety systems that guide business decisions, product growth, and prevent bad actors. If you're passionate about working with data and are eager to create solutions with significant impact, we'd love to hear from you. This role also provides the opportunity to collaborate closely with the researchers behind ChatGPT and help them train new models to deliver to users. As we continue our rapid growth, we value data-driven insights, and your contributions will play a pivotal role in our trajectory. Join us in shaping the future of OpenAI! In this role, you will: Design, build and manage our data pipelines, ensuring all user event data is seamlessly integrated into our data warehouse. Develop canonical datasets to track key product metrics including user growth, engagement, and revenue. Work collaboratively with various teams, including, Infrastructure, Data Science, Product, Marketing, Finance, and Research to understand their data needs and provide solutions. Implement robust and fault-tolerant systems for data ingestion and processing. Participate in data architecture and engineering decisions, bringing your strong experience and knowledge to bear. Ensure the security, integrity, and compliance of data according to industry and company standards. You might thrive in this role if you: Have 3+ years of experience as a data engineer and 8+ years of any software engineering experience(including data engineering). Proficiency in at least one programming language commonl
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a System Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate low-level system software for the manageability and health of OpenAI's first-party AI hardware systems. You will work across BMC, Linux, firmware interfaces, automation infra, boot and recovery, hardware diagnostics, telemetry, host and platform drivers, network software interfaces, and manufacturing and fleet readiness. A major part of this role is owning the acceptance path for partner-delivered system software: defining requirements, reviewing code and artifacts, reproducing builds, building tests, pushing fixes, and producing the evidence needed for launch decisions. This role is hands-on and high-ownership. You will write and review low-level software, debug issues across hardware and software boundaries, build infra and automation to test and manage devices in lab, guide partner deliverables, build validation evidence, and help carry platforms from bring-up through production deployment. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Design, develop, and maintain low-level firmware and system software for first-party AI hardware manageability, including BMC software, Redfish services, gNMI telemetry, firmware update and recovery flows, BIOS/UEFI interactions, platform drivers, and hardware diagnostics. Own integration and acceptance of partner and ve
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
About the Team Safety Systems works to ensure OpenAI’s most capable models can be developed and deployed responsibly. Our work spans evaluations, safeguards, red teaming, deployment decisions, and the systems that help OpenAI understand and reduce risk as models become more capable and widely used. Within Safety Systems, the Trustworthy AI team is growing its safety transparency function: a practice focused on helping external audiences understand OpenAI’s technical safety work with greater clarity, rigor, and continuity. We create and improve the public artifacts that explain how our systems are evaluated for safety, what safeguards we build, what decisions we make, and where uncertainty remains. This work includes system cards, the Deployment Safety Hub, safety-related blogs, public governance documents, and other outputs that communicate technical safety topics to external audiences. It also includes building new ways to make technical safety information easier to understand, navigate, and use—including AI-assisted workflows, data visualizations, and interactive tools that make complex technical work more legible over time. About the Role We are looking for a Safety Transparency Editor to own the editorial quality of key safety transparency artifacts and systems. This is a hands-on role for someone who can write crystal-clear, pitch-perfect explanations of the hardest and highest-stakes technical safety topics that OpenAI tackles, and who can lean into AI to build systems that help the broader organization do this work better. Your core responsibility is to shape and execute how our technical safety work is externally communicated: identifying the narrative thread, exercising judgment about which details matter, determining where additional context, explanation, or supporting evidence is needed, translating complexity without sacrificing precision, and helping external audiences understand both the safety measures we’ve taken and the uncertainties that remain. To
About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. We believe that achieving our goal requires real world deployment and iteratively updating based on what we learn. The Protection Scientist Engineer, Integrity team supports this by identifying and investigating misuses of our products – especially new types of abuse. This enables our partner teams to develop data-backed product policies and build scaled safety mitigations. Precisely understanding abuse allows us to safely enable users to build useful things with our products. About the Role Protection Science Engineering is an interdisciplinary role mixing data science, machine learning, investigation, and policy/protocol development. As a Protection Scientist Engineer within Integrity and Investigations, you will be responsible for designing and building systems to proactively identify and enforce on abuse on OpenAI’s products. This includes ensuring we have robust abuse monitoring in place for new products, sustaining monitoring for existing products, and prototyping and incubating systems of defense against our highest risk harms. You will also respond to and investigate critical escalations, especially those that are not caught by our existing safety systems. This will require expert understanding of our products and data, and involves working cross-functionally with product, policy, and engineering teams. This role is based in our London office and includes participation in an on-call rotation that will involve resolving urgent escalations outside of normal work hours. Some investigations may involve sensitive content, including sexual, violent, or otherwise-disturbing material. In this role, you will: Scope and implement abuse monitoring requirements for new product launches. Improve processes to sustain monitoring operations for existing products, including developing approaches to automate monitoring subtasks. Prototype and mature into product
About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. We believe that achieving our goal requires real world deployment and iteratively updating based on what we learn. The Protection Scientist Engineer, Integrity team supports this by identifying and investigating misuses of our products – especially new types of abuse. This enables our partner teams to develop data-backed product policies and build scaled safety mitigations. Precisely understanding abuse allows us to safely enable users to build useful things with our products. About the Role Protection Science Engineering is an interdisciplinary role mixing data science, machine learning, investigation, and policy/protocol development. As a Protection Scientist Engineer within Integrity and Investigations, you will be responsible for designing and building systems to proactively identify and enforce on abuse on OpenAI’s products. This includes ensuring we have robust abuse monitoring in place for new products, sustaining monitoring for existing products, and prototyping and incubating systems of defense against our highest risk harms. You will also respond to and investigate critical escalations, especially those that are not caught by our existing safety systems. This will require expert understanding of our products and data, and involves working cross-functionally with product, policy, and engineering teams. This role can be based in either our San Francisco, or NY office and includes participation in an on-call rotation that will involve resolving urgent escalations outside of normal work hours. Some investigations may involve sensitive content, including sexual, violent, or otherwise-disturbing material. In this role, you will: Scope and implement abuse monitoring requirements for new product launches. Improve processes to sustain monitoring operations for existing products, including developing approaches to automate monitoring subtasks. Prototyp
Get new deployment strategist lead jobs by email
Daily job updates · Unsubscribe anytime