A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and providing high-quality support for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
Jobiba hiring network
Production Operator Jobs
3,233 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current production operator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Integrations Operations Engineering (IOE) team strengthens Plaid's network with financial institutions by directly resolving integration issues to increase the reliability and quality of data, and by building new integrations to grow our financial network. We sit at the nexus of engineering, a deep understanding of Plaid's products and customers, and direct financial-institution relationships: we write and ship the code that keeps connectivity healthy, we use data to focus on the issues with the greatest customer impact, and we work directly with data partners (the financial institutions and platforms themselves) to resolve the problems that can't be fixed from our side alone. The team plays a mission-critical role in ensuring industry-leading connectivity so our customers can meet ever-expanding financial-services use cases and reach as many users as possible. What You'll Do Investigate and resolve the highest-impact integration issues by writing maintainable, tested code and deploying it to production, then monitor for regression or degradation after your changes ship. Prioritize by customer impact. The team runs a business-value-based prioritization model that automatically
About the Team OpenAI didn’t begin as a traditional company. It began as an idea: that artificial intelligence could be developed in a way that benefits everyone. As a creative team, our role is to help make sure the work behind that idea is understood, as it leaves the lab and meets the world. This company works at the frontier of intelligence. Like any frontier, it’s unfinished, constantly shifting, and still being explored. Our job is to stay close to that uncertainty and help shape how this story gets told, in a way that feels grounded and human. We do this through campaigns, launches, films, brand systems, and work that doesn’t fit neatly into any of those categories yet. This team is for people who want to help shape something from the beginning. About the Role We are seeking a Web Producer to own the execution and delivery of web projects across OpenAI.com . You’ll drive work from intake through launch, translating briefs into clear production plans, coordinating cross-functional inputs, managing timelines and dependencies, and ensuring every experience is accurate, polished, and delivered on time. You’ll partner closely with Marketing, Growth, Product, Design, Communications, and Engineering to operationalize web work in a fast-moving environment. This is a hands-on production role: you’ll build pages while owning the details, quality, communication, and follow-through that make launches run smoothly. This role is based in San Francisco. The role requires a hybrid schedule, with employees in the office Monday through Wednesday. In this role, you will: Own web projects from intake through launch, translating requests and creative briefs into structured production plans, managing timelines, dependencies, and handoffs, and proactively surfacing risks. Build and update pages using existing components, templates, and design systems; implement content, layouts, imagery, assets, and metadata accurately across single-page updates and multi-page launches. Partner wit
Chez Lyft, notre mission est de servir et de connecter. Nous nous efforçons d'y parvenir en cultivant un environnement de travail inclusif où chaque membre de l'équipe a sa place et peut s'épanouir. Lyft Urban Solutions gère les systèmes de vélos et de trottinettes en libre-service qui transportent des millions de personnes dans plus de 60 villes à travers le monde. Nous avons créé le premier système de vélos en libre-service automatisé d'Amérique du Nord et restons à la pointe de la révolution de la micromobilité. Nos systèmes sont présents partout dans le monde : de Citi Bike à New York à BIXI à Montréal, en passant par Divvy à Chicago, Capital Bikeshare à Washington D.C., et des systèmes emblématiques à Austin, Barcelone, Bogota, Boston, Buenos Aires, Columbus, Detroit, Dubaï, Londres, Madrid, Mexico, Monaco, Pittsburgh, Portland, Rio de Janeiro, San Francisco, Toronto et bien d'autres villes. Nous ne sommes pas de simples fournisseurs de technologies : nous exploitons ces systèmes, en développant les logiciels, le matériel et les opérations qui permettent à la micromobilité urbaine de fonctionner à grande échelle. Envie de contribuer à façonner le quotidien de millions de personnes dans leurs villes ? Nous recherchons un(e) ingénieur(e) logiciel Android pour rejoindre l'équipe Expérience utilisateur au sein de Lyft Urban Solutions. Votre travail contribuera directement à l'expérience numérique de nos plateformes de vélos et de trottinettes en libre-service. Vous développerez les systèmes backend qui garantissent une expérience fluide et rapide pour les programmes d'abonnement, les fonctionnalités destinées aux utilisateurs et les interactions en temps réel. Ce recrutement est urgent ; l'équipe a des objectifs ambitieux pour 2026 et recherche des ingénieurs compétents, opérationnels immédiatement. Nos ingénieurs sont des personnes brillantes, pragmatiques et capables de résoudre les problèmes rapidement, assurant ainsi un déploiement continu en production. Respon
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Urban Solutions powers the bike-sharing and scooter-sharing systems that move millions of people in over 60 cities worldwide. We built North America's first automated bike-share system and continue to lead the micromobility revolution. Our systems operate globally - from Citi Bike in New York to BIXI in Montreal, Divvy in Chicago, Capital Bikeshare in DC, and iconic systems in Austin, Barcelona, Bogota, Boston, Buenos Aires, Columbus, Detroit, Dubai, London, Madrid, Mexico City, Monaco, Pittsburgh, Portland, Rio de Janeiro, San Francisco, Toronto, and beyond. We're not just technology providers - we operate these systems, building the software, hardware, and operations that make urban micromobility work at scale. Ready to shape how millions of people move through their cities every day? We're looking for an Android Software Engineer to join the Rider Experience team within Lyft Urban Solutions. Your work will directly power the digital experiences behind our bike-share and scooter-share platforms. You'll build the backend systems that make membership programs, rider features, and real-time experiences feel seamless and fast. We're moving quickly on this hire; the team has ambitious goals for 2026 and needs strong engineers who are ready to ship. Our engineers are sharp, pragmatic problem-solvers who move fast and deploy to production continuously. Responsibilities: Engineering & Delivery Design, develop, deploy, monitor, operate, and maintain elements of our Rider experience, owning components end to end Write reliable, performant, and maintainable code, with comprehensive test coverage Reduce technical debt and contribute to ongoing improvements in tooling and code structure Analyze internal systems and processes to identify opportunities for improvement and automation Participate in code r
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This position may be a hybrid or fully remote position, as decided by your manager. If designated as hybrid, you’ll divide your time between working remotely from your home and an office location, so you should live within commuting distance. If designated as remote, you’ll be working remotely from your home and may occasionally visit a GoDaddy office to meet with your team for events or meetings. Your hiring manager can share more about this role’s hybrid or remote designation. This position is not eligible to be performed in Alaska, Mississippi, North Dakota, or the Virgin Islands. GoDaddy is not currently considering candidates for this role in California, Seattle, or NYC. Join Our Team Join a team powering secure, scalable email services for millions of customers worldwide! As part of GoDaddy's Professional Email team, you'll solve complex challenges in distributed systems, cloud infrastructure, security, and AI while modernizing critical platforms that businesses rely on every day. If you enjoy owning impactful systems, working across a diverse technology stack, and building innovative solutions at scale, you'll feel right at home here. What you'll get to do... Design, build, and maintain highly available, scalable APIs and services used by millions of customers Deploy, manage, and optimize cloud infrastructure in AWS Architect and implement modern solutions that improve performance, reliability, and security Leverage AI technologies to enhance development workflows and create innovative customer experiences Monitor, troubleshoot, and resolve complex production issues using modern observability and monitoring tools Drive continuous improvement through automation, modernization, and operational excellence C
The Applied AI team designs and builds algorithmically driven features in the Datadog app. We work across a range of applications, primarily focusing on analysis on streaming data such as anomaly detection, error outliers and faulty deployment analysis. As an Applied Scientist you will work on building models and algorithms for machine learning powered features within the Datadog platform. You will work closely with our engineering and product partners to explore, build, scale and deliver these features that we incubate within the Applied AI team. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design solutions for our different use cases. You will research and benchmark relevant algorithms to find the best fit for our use-cases Leverage machine learning algorithms and statistical techniques to build new scalable product features Develop, deploy and monitor new and existing features to production Participate in our journal club by reading and presenting the latest academic research papers to the team Explore, analyze and tell the story behind high volumes of data flowing through Datadog systems Maintain and monitor the models, services and infrastructure owned by your team Participate in your team’s on-call rotation Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering, Machine Learning or related scientific field or equivalent experience You have experience working with high-scale systems and datasets including building models, applying machine learning to real business problems, and writing production data pipelines You can explain complex ideas and algorithms to non-technical audiences You care about code simplicity and performa
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
Datadog Notebooks provide customers with a collaborative surface for ad-hoc data analysis, technical documentation, incident postmortems and runbooks. The power and flexibility of Notebooks also makes it the perfect place to integrate AI tools that can augment user workflows. Our vision is that in Notebooks users can collaborate with each other and with AI agents seamlessly. We are looking for a product-oriented Senior Software Engineer to help build the AI-assisted workflows that are becoming central to how customers use Notebooks. In this role you will work closely with product and design, and own platforms for analysis workflows and context discovery. There is a real opportunity for impact here: turning Notebooks into the tool that helps customers go from uncertainty to answer, and making that knowledge retrievable and reusable across Datadog products. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do Lead the design and delivery of AI-powered product experiences for Notebooks and adjacent surfaces Develop systems that combine deterministic product capabilities with LLM-powered experiences to deliver trustworthy and explainable customer outcomes Partner closely with Product: Work hand-in-hand with the Product Manager to translate customer problems, adoption signals, and roadmap goals into concrete technical decisions and iterations Build experiences that enable Datadog capabilities to operate within third-party AI platforms, agents, and conversational environments Provide technical leadership and mentorship while helping establish AI engineering best practices Own production systems: Build and operate reliable backend services that run in the critical path of customer deployments, and be on-call for those services Who You Are You have experience with Go,
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities—like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement b
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About The Role The Security Team is responsible for securing all things Sentry: our customers, our code, and everything in between. We are a small but growing team with broad scope, high trust, and the autonomy to tackle hard security problems with creativity and an engineering mindset. We work at a company with a strong developer culture, building a product that millions of developers genuinely love and rely on. That context shapes everything about how we operate. We take a pragmatic approach to preventing and responding to security risks. In this role not only will you build and contribute to systems which detect malicious activity, you will have the unique opportunity to implement new controls to prevent future incidents. You will work across detection and response and corporate security domains. You'll contribute to practices that keep Sentry secure as we grow: alert triage for corporate and production, detection engineering, deploying preventative controls, identity and access management, investigations and incident response, and more. You'll partner with teams across the company to prevent and respond to security incidents. You will work as a technical collaborator who prioritizes preventative controls, defense in depth, and high signal alerting practices. As Sentry expands our agentic product capabilities and development practices, you'll also find yourself at the frontier of a new set of security approaches and challenges. In this role, you will Maintain, improve, and own detection engineering systems. We own and operate our own detection stack and are building agentic triage with thoughtful security response and orc
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About The Role The Security Team is responsible for securing all things Sentry: our customers, our code, and everything in between. We are a small but growing team with broad scope, high trust, and the autonomy to tackle hard security problems with creativity and an engineering mindset. We work at a company with a strong developer culture, building a product that millions of developers genuinely love and rely on. That context shapes everything about how we operate. We take a pragmatic approach to preventing and responding to security risks. In this role not only will you build and contribute to systems which detect malicious activity, you will have the unique opportunity to implement new controls to prevent future incidents. You will work across detection and response and corporate security domains. You'll contribute to practices that keep Sentry secure as we grow: alert triage for corporate and production, detection engineering, deploying preventative controls, identity and access management, investigations and incident response, and more. You'll partner with teams across the company to prevent and respond to security incidents. You will work as a technical collaborator who prioritizes preventative controls, defense in depth, and high signal alerting practices. As Sentry expands our agentic product capabilities and development practices, you'll also find yourself at the frontier of a new set of security approaches and challenges. In this role, you will Maintain, improve, and own detection engineering systems. We own and operate our own detection stack and are building agentic triage with thoughtful security response and orc
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? North is Cohere's cutting-edge AI workspace platform, designed to transform how enterprises leverage AI. It provides a secure, customizable environment that enables organizations to deploy AI solutions while maintaining strict control over sensitive data. The North app is deployed directly to customers in highly regulated industries, delivering a seamless experience that empowers teams to work efficiently and securely at their best. We're seeking iOS and Android experts to join our close-knit team, where you'll drive exceptional impact on mission-critical projects. As a Senior Mobile Engineer, you will: Build and ship features for North, our AI workspace platform Collaborate with UI/UX designers to implement interface elements As security and privacy are paramount, you will sometimes need to re-invent the wheel, and won’t be able to use the most popular libraries or tooling Collaborate with researchers to productionize state-of-the-art models and techniques Research new technologies and best practices You may be a good fit if you: Have shipped (lots of) Android and iOS code (Kotlin, Swift) in production Built nati
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale's need to detect and respond to security events across its production and corporate environments is growing as the company scales. We're looking for a Senior Detection and Response Engineer to own detection engineering and to lead incident response when it counts, coordinating the response and driving it to resolution. This is a high-ownership role with real room to shape how detection and response works at Anyscale. You will own the detection pipeline, the response runbooks, and incident response, reporting to the Head of Security and partnering with engineering. This role is based in India. In your first year, success looks like strong detection coverage across our cloud, endpoint, and runtime telemetry, a working correlation and alerting pipeline, and incident response runbooks that have been exercised in practice. What You'll Do Own and build detection coverage across cloud, endpoint, and runtime telemetry. Own a centralized correlation and alerting capability that turns telemetry into actionable detections. Own incident response: runbooks, escalation paths, and coordination during an incident, across corporate and production environments. Drive detection of anomalous activity across the environments
Get new production operator jobs by email
Daily job updates · Unsubscribe anytime