A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure, across both cloud & on-prem environments. Site Reliability Engineers combine engineering experience and an innate drive to improve existing systems and processes, with the creativity to develop novel solutions to evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. You’ll be the experts for the environments that you operate infrastructure in, helping partner teams build & configure their software to operate reliably within. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate and participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
Jobiba hiring network
Production Operator Jobs
3,233 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current production operator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Backend Software Engineers at Palantir build software at scale to transform how organisations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Your day-to-day workflow will vary, adapting to the requirements of our users and the technical challenges that arise. One day, you may find yourself collaborating with other engineers to architect a new system that enables a novel workflow, the next you could be fine-tuning performance to enable low-latency operational outcomes. Our Product Development organisation is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product and work collaboratively to build cross functional capabilities, streamline user workflows and continuously improve our software's efficiency and reliability. We’re hiring engineers who are passionate about solving real-world problems and empowering both developers and end-users to work optimally. If you’re motivated to develop reliable, performant, and scalable systems, and to design robust APIs and primitives, this role offers the opportunity to make a sig
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments. Core Responsibilities Maintaining availability of cloud & physical Ku
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and implementation. On a daily basis, you will build end-to-end workflows, take them to production, and solve real world problems at the largest scale. You will have ample opportunity to contribute learnings from the field back to the Palantir AIP product suite. You will be on the forefront of extending Palantir's existing footprint and strategy into new markets and problem spaces opened up by Gen AI. Core Responsibilities Forward Deployed AI Engineers’ responsibilities look similar to those of a hands-on AI startup CTO: you’ll work in small teams to own delivery of high stakes projects with clients. A day’s work may include building LLM workflows on a large scale, interacting with customers to understand their needs and set their AI strategy, but the most impact will be driven by implementing solutions into the real world of our partner's organizations. Do you aspire to be an entrepreneur or an Applied AI leader? We believe Palantir is the best place — with the best colleagues — to learn how!
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and implementation. On a daily basis, you will build end-to-end workflows, take them to production, and solve real world problems at the largest scale. You will have ample opportunity to contribute learnings from the field back to the Palantir AIP product suite. You will be on the forefront of extending Palantir's existing footprint and strategy into new markets and problem spaces opened up by Gen AI. Core Responsibilities Forward Deployed AI Engineers’ responsibilities look similar to those of a hands-on AI startup CTO: you’ll work in small teams to own delivery of high stakes projects with clients. A day’s work may include building LLM workflows on a large scale, interacting with customers to understand their needs and set their AI strategy, but the most impact will be driven by implementing solutions into the real world of our partner's organisations. Do you aspire to be an entrepreneur or an Applied AI leader? We believe Palantir is the best place — with the best colleagues — to learn how!
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure, primarily across on-prem environments for the US Government. Forward Deployed Site Reliability Engineers combine engineering experience and an innate drive to improve existing systems and processes, with the creativity to develop novel solutions to evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. You’ll travel to various locations where you will be the expert for Palantir’s infrastructure, helping partner teams build & configure their hardware and network for software to operate reliably within. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate and participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s
Location Details: Remote, Canada - British Columbia or Ontario At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team We are seeking a highly creative and technically skilled Senior Motion Designer to join our creative team. In this role, you will bring static designs to life, bridging the gap between graphic design, animation, and video production. You will create high-quality motion graphics, kinetic typography, and character animations for a variety of digital platforms, including product launches, web interfaces, and brand campaigns. The ideal candidate has a strong eye for detail, composition, and pacing, and is able to translate complex ideas into visually compelling and fluid animations. This role also requires strong design skills and the ability to execute crafted designs that support motion and video projects. What you'll get to do... Collaborate with art directors, copywriters, team members, and cross-functional design partners to develop visual concepts and ensure assets are optimized for various platforms. Translate abstract ideas and scripts into style frames and storyboards, establishing the visual pacing, tone, and direction of projects before production. Create 2D (and occasionally 3D) animations, kinetic typography, visual effects (VFX), and dynamic scenes using vector illustrations, photographs, and text assets. Integrate audio, music, and voiceover with visual elements to build cohesive final deliverables. Ensure motion assets align with brand guidelines, maintain organized project files, and incorporate creative direction, client feedback, and peer feedback into iterative design improvements. Y
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available. As a Senior Site Reliability Engineer, you'll take direct ownership of production services — from initial design through day-to-day operation — while partnering with product, engineering, and security teams to build and maintain business-critical systems. In this role, you will deepen your technical expertise and grow your leadership presence by mentoring the next generation of SREs. You will also gain hands-on experience with intelligent tooling in real-world workflows. What you'll get to do... Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana Lead incident response end-to-end — performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation — validating all outputs before use Your experien
We are looking for a Lead, Data Development, AI Platform to lead the team building and operating the technical platform behind Hootsuite's Analytics MCP. This includes routing infrastructure, orchestration services, agent infrastructure, and data pipelines that enable reliable AI-driven analysis at scale. You will set technical direction for the team, strengthen engineering practices, and translate target architecture and integration standards into secure, production-grade systems that Data Analytics & AI teams can confidently build on. This is a hands-on technical leadership role. You will stay close to the code while owning delivery outcomes, engineering quality, and the team's culture of craft and accountability. You will develop strong, well-reasoned recommendations on how the platform and orchestration architecture should evolve, seek approval at the Senior Manager and Director level, and then guide the team through disciplined execution. You will also partner closely with AI Context & Integration and AI Data Architecture to ensure the platform, context layer, semantic layer, and downstream agent workflows operate as one coherent system. WHAT YOU’LL DO: Lead the design and delivery of the orchestration layer that connects the Analytics MCP platform across its most complex surfaces, including query execution across schemas and models, federated data access between the data warehouse and external source systems, and multi-step agent workflows for cross-functional business processes. Develop clear technical recommendations for platform and orchestration architecture evolution, align those recommendations with Senior Manager and Director-level direction, and guide the team through disciplined execution within the approved architecture. Own the operational reliability, scalability, quality, and observability standards for core Analytics MCP components, including routing and agent infrastructure. Guide the team to build and operate these systems to prod
We’re looking for an Intermediate Software Developer, Backend who can help us support the development organization to deliver value to customers in a reliable, efficient, and safe manner. You’ll be working in a focused team that owns one piece of the production application environment and the developer experience, you will execute on defined projects to achieve team-level goals. In line with Hootsuite's distributed workforce strategy, our flexible work arrangement allows for a hybrid model. This role is open to applicants located in Bucharest, Romania. WHAT YOU’LL DO: Write software - tools, libraries, automation, services Design and build our infrastructure platform Identify and implement new platform features Research and evaluate new technologies Refactor, rewrite or retire existing platform features Operate our developer experience and production application environments Diagnose and repair our distributed systems Perform maintenance, upgrades and migrations Control or eliminate repetitive tasks, alert noise, and business-as-usual work Enable development teams Provide executable interfaces to our infrastructure platform Provide tools and best practices to support the entire software development lifecycle Participate in a flexible on-call rotation Communicate by writing documentation, participating in meetings, and showing off your work at demos WHAT YOU’LL NEED: A degree in Computer Science or Engineering or equivalent experience working in a software engineering role An ability to write software and working knowledge of software engineering practice (Java programming language and strong working knowledge of object-oriented programming concepts) Proven experience creating stable, reliable, performing and maintainable code Familiarity with data modeling and schema design Knowledge of data structures and algorithms Open Communication: clearly conveys thoughts, both written and verbally, listening attentively and asking questions for clarification
ROLE DESCRIPTION: We’re looking for a Senior Platform Backend Developer who can help us support the development organization to deliver value to customers in a reliable, efficient, and safe manner. You’ll be working in a focused team that owns one or more pieces of the production application environment and the developer experience, you will own and deliver in service of quarterly goals on the team. ABOUT THE TEAM: This role is within our Backend Platform team. The team primarily uses Go, Scala, and PHP and has expertise in technologies such as Kafka, various AWS services, and some infrastructure-as-code tools. Your primary focus will be on developing services and tools for our product development teams as well as modernizing our existing platform. Based out of British Columbia, you will report to the Senior Manager, Software Development, DevOps. WHAT YOU’LL DO: Design and build software - tools, libraries, automation, services, and glue scripts Responsible for the reliability, security, and integrity of our large, cloud-based platform Participate in a flexible on-call rotation Lead by owning project milestones, epics or features Practice continuous improvement, contributing to culture, process, and direction in your team and across our department Develop processes and automation to eliminate repetitive tasks Design and build our infrastructure platform Identify and implement new platform features Research and evaluate new technologies Refactor, rewrite or retire existing platform features Operate our developer experience and production application environments Diagnose and repair our distributed systems Perform maintenance, upgrades, and migrations Control or eliminate repetitive tasks, alert noise, and business-as-usual work Enable development teams Provide executable interfaces to our infrastructure platform Provide tools and best practices to support the entire software development lifecycle Collaborate with others across the orga
We’re looking for an Senior Software Developer, Backend who can help us support the development organization to deliver value to customers in a reliable, efficient, and safe manner. You’ll be working in a focused team that owns one piece of the production application environment and the developer experience, you will execute on defined projects to achieve team-level goals. In line with Hootsuite's distributed workforce strategy, our flexible work arrangement allows for a hybrid model. This role is open to applicants within commutable distance to Luxembourg. WHAT YOU’LL DO: Write software - tools, libraries, automation, services Design and build our infrastructure platform Identify and implement new platform features Research and evaluate new technologies Refactor, rewrite or retire existing platform features Operate our developer experience and production application environments Diagnose and repair our distributed systems Perform maintenance, upgrades and migrations Control or eliminate repetitive tasks, alert noise, and business-as-usual work Enable development teams Provide executable interfaces to our infrastructure platform Provide tools and best practices to support the entire software development lifecycle Participate in a flexible on-call rotation Communicate by writing documentation, participating in meetings, and showing off your work at demos WHAT YOU’LL NEED: A degree in Computer Science or Engineering or equivalent experience working in a software engineering role An ability to write software and working knowledge of software engineering practice (Java programming language and strong working knowledge of object-oriented programming concepts) Proven experience creating stable, reliable, performing and maintainable code Familiarity with data modeling and schema design Knowledge of data structures and algorithms Open Communication: clearly conveys thoughts, both written and verbally, listening attentively and asking questions for clarific
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is hiring a Senior Machine Learning Operations Engineer to architect our machine learning production lifecycle. Your mission is to maintain and deploy ML models to a scalable, reliable, and secure production environment. You will design and maintain the infrastructure, automation, and monitoring systems that ensure our AI products are high-performing and cost-effective. You will report to our Director, Analytics Engineering & Data Governance and work from our Bangalore, India office. You Will: Model and Pipeline Automation Automate the deployment and retraining of ML models, from training through to production inference, by building and managing complete CI/CD/CT (Continuous Training) pipelines, adhering to MLOps best practices. Build, fine-tune, or use pre-trained LLMs, deep learning models or traditional machine learning models. Evaluate and recommend AI or ML solutions for the product using any combination of vendor solutions and/or custom-built models. Governance & Compliance Implement model versioning, lineage tracking, and auditing to ensure compliance with security and ethical standards. Performance Monitoring Continuously monitor the health and performance of production machine learning models, proactively identifying and correcting model drift, staleness, and performance degradation. Incorporate user feedback for iterative improvements and manage necessary model retraining cycles. Cross-Functional Collaboration Act as the "glue" between Data Scientists (who build models
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Technology Eva
Get new production operator jobs by email
Daily job updates · Unsubscribe anytime