Jobiba hiring network

Production Support Engineer Cloud Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current production support engineer cloud jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
14 hrs ago

Locations: South Jordan, UT Salary: $56,000 Launch Your Career in Technology Every app, website, payment, and digital service relies on technology running smoothly behind the scenes. When something goes wrong, Production Support Engineers are the people who investigate the issue, restore service, and help prevent it from happening again. If you're curious, analytical, and enjoy solving problems, this is an opportunity to build hands-on experience with cloud platforms, Linux, automation, databases, and large-scale enterprise systems from day one. What Is Production Support? Production Support Engineers keep business-critical applications running reliably in live environments. Think of it this way: Software Engineers build the platform. QA Engineers test the platform. Production Support Engineers keep the platform running when it matters most. Working at the intersection of technology and business, you'll troubleshoot issues, automate processes, and help improve the reliability and performance of systems used by thousands, or even millions, of people every day. If you enjoy solving puzzles, working under pressure, and understanding how large-scale systems work, this could be the perfect place to start your career. What You'll Do As part of a global production engineering team, you'll: Help support large-scale applications and platforms used by leading organizations around the world. Monitor business-critical applications and services to ensure high availability and performance. Investigate and resolve production incidents across applications, infrastructure, databases, and cloud environments. Analyse logs, alerts, and system metrics to identify root causes and prevent recurring issues. Partner with software engineers, infrastructure teams, and business stakeholders to improve system reliabilit

javascriptpythonjava
View job →
G
greenhouse,OneTrust
📍 BengaluruFull-time
17 hrs ago

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. Why is this a critical role at OneTrust? What is the challenge / type of challenges someone in this role will have the opportunity to take on? What impact does someone in this role have the opportunity to make within the company, for our customers and on a larger scale in the evolution of privacy and trust? An awareness of current issues affecting the industry and its technologies Create innovative, scalable, fault-tolerant software solutions for our clients and customer base Expand existing software to meet the changing needs of our key demographics What does this person do each day/each week? Describe a true to life day in the life for someone in this role. What do they do and how do they do it? Goal is to paint a real, genuine picture so candidates can see themselves in the role. Support production customers by monitoring and maintaining our cloud application & cloud infrastructure hosting it Build scripts for operational automation and incident response Handle processes surrounding cloud application deployment for our agile release Work with the monitoring, tuning, maintenanc

pythonjavasql
View job →
G
greenhouse,WPP
📍 LisbonFull-time
17 hrs ago

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: At WPP, technology is at the heart of everything we do, and it is the Technology Operations teams mission, as part of Enterprise Technology , to enable our stakeholders to collaborate, create and thrive. Enterprise Technology is undergoing a significant transformation to modernise ways of working, shift to cloud and micro-service-based architectures, drive automation, digitise colleague and client experiences and deliver insight from WPP’s petabytes of data. This role will carry out the effective and efficient everyday technology operations for WPP ET. A trusted pair of hands to deal with level 1 and 2 issues as they present to the IT Service Desk and a trusted resource for Infrastructure and Management personnel to assist with project work when needed. The role will report into the Enterprise Technology Operations Lead and work closely with other teams within Enterprise Technology. What you'll be doing: Deliver world class, on-site support services to WPP employees, agencies, and visiting clients, operating within predefined structur

G
17 hrs ago

A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO As Database Support Engineer, you’ll support various critical database platforms across Development, QA, UAT, and Production environments. The role partners closely with application teams, application support, and database engineers and operates within a Follow‑the‑Sun model to ensure availability, performance, and reliability of database services. Key responsibilities include: • Provide operational support for enterprise database platforms in both on-prem private cloud and public cloud • Monitor database health, capacity, performance, and availability, and respond to alerts, diagnose issues, and perform timely remediation • Perform routine maintenance activities (patching, upgrades, housekeeping etc) • Troubleshoot database‑related incidents and collaborate on root cause analysis • Work closely with application owners, application support teams, and DB Engineers • Provide guidance on database best practices and operational standards • Participate in cross‑team problem resolution and continuous improvement initiatives • Contribute to design, implementation and testing of automation and self service capabilities of DB platforms • Drive continuous improvement, identifying opportunities to reduce toil and increase platform efficiency. • Participate in a Follow‑the‑Sun operating model, including shift‑based coverage and handoffs WHAT’S REQUIRED • Bachelor’s degr

pythonsqlpostgresql
View job →
G
greenhouse,WPP
📍 ChennaiFull-time
17 hrs ago

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: You will play a crucial role in solution delivery and day-to-day technology management of the global enterprise automation and AI program. Reporting to the Support & Ops Manager and supporting AI product, platform and delivery teams worldwide, you will ensure our AI projects are well managed and supported to deliver quantifiable returns. Our team is chartered to develop and deliver custom AI products and enhance global cloud apps products or AI projects delivered by other delivery teams from around our global firm. Our support apparatus gives those teams confidence and freedom to hand over their existing successes and continue innovating. Are you a proactive and detail-oriented individual with a passion for technology and a strong desire to learn and grow in the field of automation & AI? We are looking for a L2/L3 Automation Support Engineer to join our dynamic team. You'll be responsible for the monitoring, tria

sqlazuregit
View job →
D
Datadog
📍 MassachusettsFull-timeFrom $244K/yr
1mo ago

Datadog’s Cloud Networks team designs, builds, and maintains the production network infrastructure that powers everything built on top of our platform across AWS, GCP, Azure, and beyond. In this role, you’ll set technical direction for how we scale our multi-region, multi-cloud network footprint while keeping reliability and performance high. You’ll partner closely with internal teams and Cloud Service Providers to troubleshoot complex connectivity issues, integrate new networking capabilities, and improve the foundations our engineers and customers rely on. This is a high-impact opportunity to drive meaningful improvements in scale, resiliency, and cost efficiency. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design, build, and operate cloud network infrastructure across AWS, GCP, Azure, and Neoclouds in a multi-region environment. Own connectivity between clouds, customers, and developers—ensuring scalable, secure, and reliable network paths. Set clear technical direction for expanding data centers and evolving the network while maintaining stability and performance. Improve cross-site and cross-region connectivity patterns to support Datadog’s growing platform needs. Lead deep investigations into latency, packet loss, and connectivity failures – from pcap and path analysis through to escalations with cloud providers that may originate from customer support Identify and deliver network-related efficiency and cost-saving opportunities that positively impact business health. Who You Are: You have deep networking expertise. You understand BGP, route policies, path selection, prefix advertisement, and what breaks in large-scale networking. You have substantial experience designing, building, and evolving large-scale Software-Defined Networks—inclu

awsazuregcp
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. As a Senior Technical Support Engineer, you will be a trusted advisor and technical resource for our customers in the EMEA region. This is a hands-on role for someone who thrives in dynamic environments, loves troubleshooting complex technical issues, and is passionate about delivering exceptional support experiences. You’ll be responsible for resolving high-impact technical issues, driving customer success, and collaborating closely wit

pythonsqlaws
View job →
G
Godaddy
📍 British ColumbiaFull-timeFrom C$107K/yr
25 days ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operati

pythonawsci/cd
View job →

Location Details: At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operat

pythonawsci/cd
View job →

Here’s a summary of the role: Build cloud software that matters, grow your technical depth, and use modern AI tooling to do your best work. This is a hands-on engineering role for someone who enjoys solving product problems, writing clean code, and helping services run reliably at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS , and modern engineering practices. You’ll be part of a collaborative product engineering team where you can own features, contribute to design discussions, support production systems, and keep growing across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Design, build, test, and improve backend services and APIs using Node.js, TypeScript, and AWS . Take ownership of well-defined features from planning through release, including code quality, deployment, and production support . Work closely with product managers, designers, and other engineers to turn requirements into practical, reliable solutions. Contribute to technical design conversations, code reviews, and engineering standards that keep the team moving well. Use AI tools to speed up research, coding, debugging, testing, and documentation, while checking outputs carefully and applying sound judgment. Help keep systems secure, observable, and maintainable by improving monitoring, reliability, and day-to-day development practices. These are the essentials you’ll need to get an interview: 3 to 5 years of professional software engineering experience building production applications in an agile environment. Strong backend development skills with Node.js and TypeScript, including experience building APIs or microservices. Experience with React or Angular in a product engineering environment. Hands-on experience with

typescriptreactnode.js
View job →
L
1mo ago

At Lyft, community is what we are and it’s what we do. It’s what makes us different. To create the best ride for all, we start in our own community by creating an open, inclusive, and diverse organization where all team members are recognized for what they bring. Our core corporate functions (Finance, Supply Chain) are critical to Lyft’s success, and the health of our systems and processes is essential to daily operations and long-term growth. This position requires strong expertise in Oracle Procurement Cloud and Oracle Accounts Payables Cloud, with hands-on experience in solution design, configuration, and production support. You will support and continuously improve Lyft’s corporate systems by managing incidents, enabling automation, implementing monitoring, configuring applications, and driving end-to-end Source-to-Pay lifecycle excellence. To effectively support our business stakeholders, candidates for this role must be proactive, detail oriented, highly subject oriented, analytical, customer obsessed and possess the ability to execute the following skills. Your Key Responsibilities: Serve as a skilled Oracle Fusion Source-to-Pay (STP) Functional professional, acting as a strategic bridge between business and technical teams to support and enhance ERP Financials, Procurement, Payables, Supplier Management, and related modules within Oracle Fusion Cloud Applications Source to Pay Workstream. Apply functional expertise alongside hands-on technical proficiency in integrations, reporting, and issue resolution, delivering results in a dynamic, client-facing environment. Support end-to-end STP processes, including Supplier Setup & Qualification, Procurement & Sourcing, Purchase Order Management, Self-Service Purchasing, Payables, Expenses, and Supplier Portal. Bring 5–10 years of hands-on implementation experience across Oracle Financials and Supply Chain modules, with strong exposure to Accounts Payables, Procurement and STP operations. Experience in zip, a

sqlagileai
View job →

Here’s a summary of the role: Build software that matters, take real technical ownership, and use modern AI tooling to do your best work. This is a hands-on senior engineering role for someone who enjoys solving complex product problems, shaping robust solutions, and helping teams deliver reliable services at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS, and modern engineering practices. You’ll play a leading role within a collaborative product engineering team, owning complex features end to end, contributing to design and architectural decisions, supporting production systems, and helping raise the bar across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Own and deliver complex backend services and APIs using Node.js, TypeScript, and AWS , from technical design through release and production support. Contribute to design and architecture discussions, making pragmatic decisions that balance delivery speed, maintainability, scalability, and security. Mentor and support less experienced engineers through code reviews, pairing, technical guidance, and day-to-day collaboration. Work closely with product managers, designers, and engineers across the team to turn requirements into practical, reliable solutions. Use AI tools to accelerate coding, debugging, testing, research, and documentation, while validating outputs carefully and applying sound judgment. Strengthen service reliability, observability, and engineering quality by improving monitoring, incident response, testing, and development practices. These are the essentials you’ll need to get an interview: 5 to 8 years of professional software engineering experience delivering production systems in an agile environment. Strong backend development s

typescriptreactnode.js
View job →
GH
greenhouse,Cohere Health
📍 HyderabadFull-time₹2K – ₹2K/yr
18 hrs ago

Opportunity Overview: We are seeking a Lead Software Engineer to join our Integrations team. In this role, you will be designing, developing, and scaling highly available healthcare integration systems supporting prior authorization workflows across providers, payers, and delegated entities. You'll direct a fast-paced, autonomous,agile team of software engineers in the design, development, and operational support of a growing enterprise integration platform. This is an opportunity to drive technical excellence at the intersection of healthcare interoperability and modern distributed systems. What you’ll do: Technical Leadership: Provide technical leadership across architecture, system design, platform scalability, reliability, and operational excellence. Platform Engineering: Design and build scalable, resilient, and high-performing systems that support critical business workflows and enterprise integrations. Integration Solutions: Lead the development and maintenance of secure integrations with internal and external platforms, partners, and third-party systems. Cloud & Automation: Drive cloud infrastructure, deployment automation, and software delivery practices that enable reliable and efficient releases. Distributed Systems: Design and support event-driven and distributed architectures that enable scalable and fault-tolerant processing. Operational Excellence: Establish monitoring, observability, and incident response practices to ensure system reliability, performance, and availability. Quality Engineering: Champion automated testing, quality assurance, and engineering best practices throughout the software development lifecycle. Production Support: Lead the resolution of complex production issues and drive continuous improvement in platform stability and operational efficiency. Cross-Functional Collaboration: Partner with product, operations, data, security, and business stakeholders to deliver solutions aligned with organizational goals. Agile Delive

javaawsdocker
View job →
GH
18 hrs ago

Opportunity Overview: Cohere Health is at the forefront of leveraging artificial intelligence (AI) to transform prior authorization, shifting from transaction-based processes to creating elevated care journeys. By integrating AI into healthcare decisioning, we aim to improve efficiency, reduce costs, and ultimately enhance patient care. This is a unique opportunity to join our rapidly growing Clinical Intelligence team, where you’ll work on building impactful healthcare technologies that streamline and optimize the prior authorization process using AI/ML and Advanced Rules Engines. What you’ll do: Contribute to the functional & technical roadmap for the team, working alongside cross-functional teams to deliver high-impact deterministic rules-based and AI-driven healthcare automation solutions. Actively support the technical design process, bringing your expertise and analysis to help make data-driven decisions Continuously discover, understand, and implement new technologies & services to maximize development efficiency Contribute heavily to feature design, development, testing, and delivery of our cloud platform and web applications Support all parts of our platform from the database to the frontend Drive daily engineering release once 2-4 times a month Perform production support duties 1-2 times a year. ISMS roles and responsibilities: Good knowledge of Information practices. Assist the manager in all the information security activities implementation and maintenance process. Ensuring the team and imparted with Competence related to Information security Responsible for implementation of security policies and procedures and report any issues to the Information Security Manager. What we looking for: You have experience working on software development teams, building and deploying full stack web applications You are passionate about building quality products and want to own product development end-to-end, with excellent design and development stan

javascripttypescriptpython
View job →
GH
18 hrs ago

Opportunity Overview: Cohere Health is at the forefront of leveraging artificial intelligence (AI) to transform prior authorization, shifting from transaction-based processes to creating elevated care journeys. By integrating AI into healthcare decisioning, we aim to improve efficiency, reduce costs, and ultimately enhance patient care. This is a unique opportunity to join our rapidly growing Clinical Intelligence team, where you’ll work on building impactful healthcare technologies that streamline and optimize the prior authorization process using AI/ML and Advanced Rules Engines. What you’ll do: Contribute to the functional & technical roadmap for the team, working alongside cross-functional teams to deliver high-impact deterministic rules-based and AI-driven healthcare automation solutions. Actively support the technical design process, bringing your expertise and analysis to help make data-driven decisions Continuously discover, understand, and implement new technologies & services to maximize development efficiency Contribute heavily to feature design, development, testing, and delivery of our cloud platform and web applications Support all parts of our platform from the database to the frontend Drive daily engineering release once 2-4 times a month Perform production support duties 1-2 times a year. ISMS roles and responsibilities: Good knowledge of Information practices. Assist the manager in all the information security activities implementation and maintenance process. Ensuring the team and imparted with Competence related to Information security Responsible for implementation of security policies and procedures and report any issues to the Information Security Manager. What we looking for: You have experience working on software development teams, building and deploying full stack web applications You are passionate about building quality products and want to own product development end-to-end, with excellent design and development stan

javascripttypescriptpython
View job →
🔔

Get new production support engineer cloud jobs by email

Daily job updates · Unsubscribe anytime