Jobiba hiring network

Production Cleaning Specialist Jobs

3,233 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current production cleaning specialist jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

awsazurerest
View job →
O
1mo ago

About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p

pythonawsci/cd
View job →
O
OpenAI
📍 San Francisco• Full-time• $342K – $445K/yr
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly-Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring-up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross-functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring-up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross-functional judgment, and the ability to drive clarity in ambiguous, fast-moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring-up and validation through production deployment, operati

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Strategic Sourcing & Procurement function plays a critical role in enabling OpenAI to deliver impact across research, product development, technology infrastructure, and services by helping the company scale responsibly, securely, and with strong commercial discipline. Our work sits at the intersection of innovation and execution. We partner closely with teams across OpenAI to translate rapidly evolving business needs into scalable, compliant, and economically sound external partnerships. As OpenAI continues to grow at pace, services sourcing is becoming increasingly strategic across the company. Every business unit relies on external service providers in different ways — to extend capacity, access specialized expertise, support operations, and accelerate execution. Done well, Procurement becomes a source of trust and momentum, helping OpenAI move faster with the right partners, stronger commercial outcomes, and the right level of protection. About the Role We are seeking an experienced Strategic Sourcing (GTM) Leader to lead strategic sourcing and commercial enablement for OpenAI’s Go-to-Market organization across B2B and B2C channels. You will manage substantial and rapidly growing spend while shaping sourcing strategies and scalable commercial pathways across Media, Creative, Production, Influencer, Agency, Sponsorships, Analytics, Communications, and Event suppliers in support of high-impact global initiatives. You’ll help evolve our GTM procurement function from reactive deal support into a speed-enabling, scalable commercial engine that delivers cost efficiency, launch readiness, and strong governance in a fast-moving environment. In this role, you will: Develop and execute sourcing strategies across GTM, Brand, Global Affairs, Events, Growth, and Partnership activities—spanning both B2B and B2C channels—that align with our mission and b

reactawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The AI Architect team partners with organizations to turn OpenAI's most capable models into meaningful, real-world impact. We work with customers across industries and digital-native businesses to identify where AI can create value, design secure and scalable solutions, and help those solutions move from early exploration into sustained production adoption. The team brings together technical strategy, customer partnership, and practical deployment expertise, working closely with Sales, Product, Engineering, Research, and specialist delivery teams. About the Role As an AI Architect, you will be the senior technical owner for a named portfolio of customers and the primary technical counterpart to their leadership teams. You will act as the “CTO of your book of business”, shaping each customer's AI strategy and guiding their journey from pre-sales discovery and solution evaluation through deployment, adoption, and measurable business impact. You will own the technical account plan across ChatGPT Enterprise, the OpenAI API, Codex, and other agentic AI solutions. In partnership with the Account Director, you will translate business priorities into a focused use-case portfolio, an actionable adoption roadmap, and a clear path to durable customer value and growth. The Account Director owns commercial strategy; you own the technical strategy, customer journey, and path to production value. You will remain accountable for the technical outcome while bringing in the right specialists across deployment, implementation, enablement, security, product, and partners to provide deeper expertise and execute work where needed. This role calls for strong industry fluency, sound architectural judgment, and the ability to move confidently between executive strategy and hands-on technical conversations. In this role, you will: Serve as the primary technical advisor and long-term technical relationship owner for a named portfolio of existing customers and pre-sales prospect

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

AI Systems Engineer - Codex Core Agents About The Team The Codex Core Agents team builds the agent harness that turns model capability into real-world action. We own the systems around the model: prompting and interpreting model outputs, executing actions safely in real environments, and feeding production experience back into better models and better agent behavior. This team sits close to research and works across the stack: harness, model interaction, inference, sandboxed execution, orchestration, evals, production reliability, and the performance envelope around tokens, latency, cost, capacity, and quality. The harness is open source and increasingly part of how models are trained and evaluated, making this one of the highest-leverage layers in Codex. About The Role We’re looking for engineers to build the AI systems that make Codex agents dependable in production. The ideal candidate is an agent-systems builder: hands-on across low-level systems and ML workflows, able to debug Codex behavior end to end across the harness, model behavior, inference/runtime stack, GPU fleet, and product surface. You’ll work with research, infrastructure, and product to design agent harness capabilities, run experiments and ablations across the model + system prompt + harness stack, build frameworks for assessing production agent performance, and turn messy failures into durable improvements. What You’ll Do Design and build the core agent harness and execution loop that lets Codex agents interpret model outputs, use tools, execute code, and complete long-horizon tasks safely. Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents operating in real development environments. Develop evaluation, experimentation, and debugging systems that distinguish harness issues, model behavior, inference/runtime issues, and product failures. Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to

pythonawsrest
View job →
O
1mo ago

About the Team Our London-based team builds the backend systems that help ChatGPT scale reliably. We work on infrastructure close to the product, partnering with engineering teams to improve the performance, resilience, and operability of critical user-facing systems. Our work combines backend software engineering with distributed systems and production reliability. We build shared capabilities, improve high-traffic workflows, and make it easier to introduce new product functionality without compromising performance or availability. About the Role This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability. You’ll also own how those systems behave in production: how changes are rolled out, how issues are detected and diagnosed, and how recurring operational problems can be addressed through better software and system design. This is a strong fit for backend engineers who enjoy complex systems problems and want a direct connection between the infrastructure they build and the experience of ChatGPT users. In this role, you will: Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences. Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely. Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow. Build and improve systems for asynchronous processing and other large-scale backend workloads. Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on-call, incident response, and root-cause analysis, and turn operational lea

awsrestai
View job →
O
1mo ago

About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and

awskubernetesrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Senior Mechanical Engineer to lead the design, integration, and sustaining engineering of mechanical subsystems in robotic platforms. You will work closely with experienced engineers and cross-functional partners to set functional requirements, develop and iterate on hardware to meet program expectations. This role is aimed at candidates with strong fundamentals in mechanical system design — including tolerance, alignment, load paths, wear, and failure modes — and robotics. You will contribute to real hardware programs moving from prototype through early production. You will lead both new subsystem development and ongoing improvements to existing systems based on testing, field performance, and manufacturing feedback. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Lead the design and iteration of mechanical subsystems, including structures, mechanisms, and actuators. Create and maintain CAD models, assemblies, and drawings with appropriate tolerancing and documentation. Build and test prototypes, supporting debugging of mechanical issues such as fit, alignment, friction, and wear. Assist in developing test methods and executing validation to evaluate performance, durability, and failure modes. Work with cross-functional teams to integrate mechanical components with sensors, actuators, and control systems. Support transition of designs from prototype to manufacturable assemblies, incorporating DFM and DFA considerations. Collaborate with manufacturing partners

awsrestai
View job →
O
1mo ago

About the Team: GTM Innovation is a product engineering team with a charter to automate 100% of digital knowledge work in OpenAI's GTM, so sellers spend more time directly with customers. AGI-level reasoning doesn’t mean organization-level transformation “just works” out of the box; orgs must be redesigned around abundant intelligence and persistent virtual coworkers. Our team builds and scales a fleet of virtual coworkers that operate as full-time members of the account team, and redefines how our human-first revenue organization interacts with their agentic teammates. About the Role We’re looking for backend software engineers with a product mindset to join the GTM Innovation team. You’ll help OpenAI meet the world at scale. You’ll partner closely with go-to-market teams to understand their workflows, identify leverage points, and ship novel solutions using OpenAI’s API platform. You’ll move quickly from prototype to production, and your work will directly shape how customers experience our technology in the field. This role is ideal for engineers who want to be close to users, own end-to-end outcomes, and help define entirely new categories of enterprise software. In this role, you will: Build high-impact applications and tools that accelerate OpenAI’s go-to-market efforts Work across the full product lifecycle for GTM: prototype, iterate, ship, and maintain Embed with Sales, Technical Success, and Revenue Operations to identify user needs and build for them Apply OpenAI’s models in novel ways to solve real-world customer and internal workflow problems Translate learnings into feedback for Applied and Research teams to inform product development You’ll thrive in this role if you: Have 4+ years of experience as a software/ML/product engineer working on user-facing systems Former founder, or early engineer at a startup who built a product from scratch is a plus Are fluent in Python or JavaScript and comfortable building full-stack applications Have built or prototy

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role On the Accelerators team, you will help OpenAI evaluate and bring up new compute platforms that can support large-scale AI training and inference. Your work will range from prototyping system software on new accelerators to enabling performance optimizations across our AI workloads. You’ll work across the stack, collaborating with both hardware and software aspects - working on kernels, sharding strategies, scaling across distributed systems, and performance modeling. You'll help adapt OpenAI's software stack to non-traditional hardware and drive efficiency improvements in core AI workloads. This is not a compiler-focused role, rather bridging ML algorithms with system performance - especially at scale. In this role, you will: Prototype and enable OpenAI's AI software stack on new, exploratory accelerator platforms. Optimize large-scale model performance (LLMs, recommender systems, distributed AI workloads) for diverse hardware environments. Develop kernels, sharding mechanisms, and system scaling strategies tailored to emerging accelerators. Collaborate on optimizations at the model code level (e.g. PyTorch) and below to enhance performance on non-traditional hardware. Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization. Work with hardware teams and vendors to evaluate alternatives to existing platforms and adapt the software stack to their architectures. Contribute to runtime improvements, compute/communication over

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Mechanical Engineer to design and prototype novel tactile, proprioceptive and force sensor solutions. You will work closely with experienced engineers and cross-functional partners to set functional requirements, develop and iterate on hardware to meet program expectations. This role is aimed at candidates with strong fundamentals in mechanical system design and robotic sensing including tactile sensors, force sensing and position sensing. You will contribute to real hardware programs moving from prototype through early production. You will lead both new sensor subsystem development and ongoing improvements to existing systems based on testing, field performance, and manufacturing feedback. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Lead the design and iteration of mechanical subsystems, including tactile sensors, load cells, force torque sensors and joint encoders. Create and maintain CAD models, assemblies, and drawings with appropriate tolerancing and documentation. Build and test prototypes, supporting debugging of mechanical issues such as fit, alignment, friction, and wear. Assist in developing test methods and executing validation to evaluate performance, durability, calibration and failure modes. Work with cross-functional teams to integrate mechanical components with sensors, actuators, and control systems. Support transition of designs from prototype to manufacturable assemblies, incorporating DFM and DFA considerations. Collaborate with manufact

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team: GTM Innovation is a product engineering team with a charter to automate 100% of digital knowledge work in OpenAI's GTM, so sellers spend more time directly with customers. AGI-level reasoning doesn’t mean organization-level transformation “just works” out of the box; orgs must be redesigned around abundant intelligence and persistent virtual coworkers. Our team builds and scales a fleet of virtual coworkers that operate as full-time members of the account team, and redefines how our human-first revenue organization interacts with their agentic teammates. About the Role We’re looking for product mindset software engineers to join the GTM Innovation team. As a product engineer on this team, you’ll help OpenAI meet the world at scale. You’ll partner closely with go-to-market teams to understand their workflows, identify leverage points, and ship novel solutions using OpenAI’s API platform. You’ll move quickly from prototype to production, and your work will directly shape how customers experience our technology in the field. This role is ideal for engineers who want to be close to users, own end-to-end outcomes, and help define entirely new categories of enterprise software. In this role, you will: Build high-impact applications and tools that accelerate OpenAI’s go-to-market efforts Work across the full product lifecycle for GTM: prototype, iterate, ship, and maintain Embed with Sales, Technical Success, and Revenue Operations to identify user needs and build for them Apply OpenAI’s models in novel ways to solve real-world customer and internal workflow problems Translate learnings into feedback for Applied and Research teams to inform product development You’ll thrive in this role if you: Have 4+ years of experience as a software/ML/product engineer working on user-facing systems Former founder, or early engineer at a startup who built a product from scratch is a plus Are fluent in Python or JavaScript and comfortable building full-stack applications

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b

pythonawsrest
View job →
🔔

Get new production cleaning specialist jobs by email

Daily job updates · Unsubscribe anytime