Here’s a summary of the role: Build cloud software that matters, grow your technical depth, and use modern AI tooling to do your best work. This is a hands-on engineering role for someone who enjoys solving product problems, writing clean code, and helping services run reliably at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS , and modern engineering practices. You’ll be part of a collaborative product engineering team where you can own features, contribute to design discussions, support production systems, and keep growing across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Design, build, test, and improve backend services and APIs using Node.js, TypeScript, and AWS . Take ownership of well-defined features from planning through release, including code quality, deployment, and production support . Work closely with product managers, designers, and other engineers to turn requirements into practical, reliable solutions. Contribute to technical design conversations, code reviews, and engineering standards that keep the team moving well. Use AI tools to speed up research, coding, debugging, testing, and documentation, while checking outputs carefully and applying sound judgment. Help keep systems secure, observable, and maintainable by improving monitoring, reliability, and day-to-day development practices. These are the essentials you’ll need to get an interview: 3 to 5 years of professional software engineering experience building production applications in an agile environment. Strong backend development skills with Node.js and TypeScript, including experience building APIs or microservices. Experience with React or Angular in a product engineering environment. Hands-on experience with
Jobs in United Kingdom
Production Support Sre Analyst in London
52 active opportunities · Updated October 2026
Showing
15 jobs
Explore current production support sre analyst jobs in London. Filter by work mode, employment type, experience, department, date posted and distance.
Here’s a summary of the role: Build software that matters, take real technical ownership, and use modern AI tooling to do your best work. This is a hands-on senior engineering role for someone who enjoys solving complex product problems, shaping robust solutions, and helping teams deliver reliable services at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS, and modern engineering practices. You’ll play a leading role within a collaborative product engineering team, owning complex features end to end, contributing to design and architectural decisions, supporting production systems, and helping raise the bar across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Own and deliver complex backend services and APIs using Node.js, TypeScript, and AWS , from technical design through release and production support. Contribute to design and architecture discussions, making pragmatic decisions that balance delivery speed, maintainability, scalability, and security. Mentor and support less experienced engineers through code reviews, pairing, technical guidance, and day-to-day collaboration. Work closely with product managers, designers, and engineers across the team to turn requirements into practical, reliable solutions. Use AI tools to accelerate coding, debugging, testing, research, and documentation, while validating outputs carefully and applying sound judgment. Strengthen service reliability, observability, and engineering quality by improving monitoring, incident response, testing, and development practices. These are the essentials you’ll need to get an interview: 5 to 8 years of professional software engineering experience delivering production systems in an agile environment. Strong backend development s
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Role: Senior Technical Support Specialist (Enterprise Agentic AI) Company: Ema Unlimited Inc. Location: London Employment Type: Full-time/Remote 1. About Ema (Why Ema) Ema is building the world’s first Universal AI Employee — a production-grade agentic AI platform that automates real enterprise workflows across HR, IT, Finance, and Operations. Ema’s customers do not run demos. They replace mission-critical, manual business processes with agentic AI systems that operate across multiple SaaS tools, APIs, and human-in-the-loop workflows. In this world, support is not reactive . Support is production reliability, trust preservation, and system learning . At Ema, Senior Technical Support Specialists are operators of live AI systems , not ticket handlers. 2. Role Overview The Senior Support Engineer owns the health, reliability, and trustworthiness of Ema’s deployed agentic AI systems in production. This role sits at the intersection of: AI behavior Workflow orchestration Enterprise integrations Customer trust Engineering feedback loops This is: ❌ Not L1 / call-center support, ❌ Not a “just escalate to engineering” role, ❌ Not reactive firefighting only This is : A senior technical escal
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: WPP is embarking on a major 3-year transformation to simplify, modernise, and unify its technology platforms across the group. This change is critical to: Enable cost-effective transformation and operational efficiency Accelerate adoption of AI, data, and digital platforms Ensure consistent delivery of value and benefits across all agencies and markets Minimise disruption and maximise engagement during large-scale change Workday is a cornerstone of this transformation, providing a single, modern system of record for people data and for driving efficiency through our operations and finance functions. By standardising core processes, improving data quality and enabling better insights, Workday supports consistent ways of working, stronger governance and future‑ready capabilities, while creating the foundation for scalable change, automation and AI‑enabled decision making across WPP. We are looking for a Workday Global HCM Support Lead with strong experience in leading a global technology support function. The purpose of the role is to deliver a high-quality support service to our key s
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond to operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will design and develop systems that support the model lifecycle in production, including deployment orchestration, configuration management, operational automation, reliability, and capacity management. You will help transform complex operational processes into scalable platform capabilities that enable teams across OpenAI to deploy and manage models with greater confidence and less manual effort. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build and evolve the platform used to deploy, configure, and manage models across ChatGPT. Develop systems for deployment orchestration, model rollouts, operational visibility, and production readiness. Create abstractions and tooling that simplify complex infrastructure and improve the developer experience. Automate operational workflows, including incident detection, diagnosis, mitigation, and recovery. Improve the reliability, scalability, and efficiency of model deployments and the infrastructure that supports them. Build systems that support capacity planning, resource allocation, and infrastructure utilization. Partner with research, infrastructure, and product engineering teams to identify common chal
About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. About Data & Technology Solutions WPP’s Data & Technology Solutions is the WPP’s unified global data products and technology team. We work with our agencies and clients to build data-driven solutions and technology product to power marketing transformation. WPP Open is our AI platform for marketing, it connects marketing professionals, data, tools and AI in a single place. WPP Open is the simplest, safest and fastest way to realize the benefits of scaled AI – delivering better-informed creative ideas faster, at scale and at lower cost. We’re endlessly curious and our team of thinkers, builders, creators and problem solvers are over 2,000 strong, across 20 markets around the world. WHO WE ARE LOOKING FOR We are seeking a strategic and commercially driven leader to define and execute WPP’s data partnership strategy within EMEA. This role will be critical in sourcing, integrating, and scaling data partnerships to support local market growth whi
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: We are seeking a commercially minded Financial Controller to join WPP Media. This role will be responsible for leading and managing the month-end, quarter-end, and year-end close processes, ensuring all deadlines are met and reporting is delivered accurately and efficiently. The successful candidate will maintain robust financial controls, drive continuous process improvements, and support the delivery of a streamlined and effective reporting function. This function provides timely and accurate monthly reporting to WPP, including P&L and balance sheet submissions, as well as high-quality management information to support business decision-making. What you'll be doing: Financial Reporting & Month-End Close Lead the month-end, quarter-end, half-year, and year-end close processes, ensuring accurate and timely financial reporting. Prepare and review monthly management accounts, including profit and loss statements, balance sheets, and supporting commentary. Deliver insightful variance analysis, identifying key business drivers and providing recommendations to stakeholders. Overse
Agents have changed the game for software delivery and efficacy. Diligent is the leading GRC platform in the world, and we are racing ahead to take the agents show on the road and work with the customers where they work . The FDE function will lead the change on how we embed AI agents into some of the world’s most complex governance, risk and compliance environments. This is not a support or consultancy role. It is a builder role, for someone who is equally comfortable reading a failing agent trace, running a discovery workshop with a bank’s internal audit team, and translating what they find into a production-grade agentic solution. You will be building AI agents for GRC professionals , not assistants that surface suggestions, but agents that own complex, multi-step workflows end to end . Agents that customers can hand a task to and trust it will come back done. Closing the gap between a promising prototype and something a company Board and ELT depends on is a completely . Here’s a breakdown of what you’ll do Embed directly with major enterprise customers (global banks, regulated corporates) across EU and US ; sitting wi th internal audit teams, risk functions, compliance and governance professionals to understand their real workflows and devise agentic solutions to intelligently automate them creating tremendous efficacy and efficiencies for our customers. Run agent-focused discovery workshops, rapidly prototype agentic solutions, and test them with practitioners; distinguishing between workflows that need an agent and those that need a button. Source, integrate, and move data between enterprise systems as part of live customer implementations — understanding the real data landscape customers operate in and building reliable pipelines to support it . Take agents from prototype t
About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation opportunities that reduce repetitive work while preserving human review, judgment, and accountability. Own production systems through testing, observability, incident response, documentation, and continuous reliability improvements. Hel
About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and
At Rockstar Games, we create world-class entertainment experiences. Become part of a team working on some of the most rewarding, large-scale creative projects to be found in any entertainment medium - all within an inclusive, highly-motivated environment where you can learn and collaborate with some of the most talented people in the industry. Rockstar is on the lookout for a talented Software Engineer who possesses a strong interest in all the low-level technology that makes a modern video game tick to support the Cfx.re creator platforms, including FiveM and RedM. As a member of our team, you will need a critical and creative eye capable of putting forth innovative solutions to complex problems. If you like to understand how things really work “under the hood” of your favorite games, we’d love to hear from you. This is a full-time, permanent and in-office position based in Rockstar’s unique game development studio in the heart of London. WHAT WE DO The Rockstar Creator Platform Team deliver a technology platform that enables players to experience community created content on fully customized dedicated servers where creators can develop their own game modes and other modifications in a variety of scripting languages. We create technology, tools, and solutions to enhance the creator experience and empower our community to create and share any experience imaginable. RESPONSIBILITIES Maintain and improve existing and new codebases, ensuring high standards of quality, stability, and efficiency in collaboration with cross-functional teams. Support software release processes across multiple branches, coordinating with production and engineering teams to ensure readiness and a high standard of quality. Help identify, prioritize, and resolve critical issues, coordinating timely fixes while maintaining overall system stability. Drive release planning, deployment, and rollback procedures. Maintain and enhance build, test, and release automation
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and providing high-quality support for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. The role “The ideal forward deployed builder is a composite. Not a software engineer, not a product manager, not a solutions architect or a salesperson — but able to draw on any of these to solve the customer’s problems, in the customer’s building, on the customer’s timeline, against the customer’s data.” Ema is building the Universal Agentic Control Plane — the layer where enterprises turn messy, human-run business processes into AI Employees that actually do the work. The platform is powerful. The bottleneck is people who can walk into a customer’s world, see what should be automated, and make it real. That’s this role. We’re looking for a Forward Deployed Builder to embed inside our customers and partners, find the highest-value opportunities for agentic automation, solution them, build the demo, win the deal, and then stay to implement it in production. You go in alone, Rambo in the jungle, armed with your wits, your skills, and the most powerful agentic platform on the market. We’ll train you hard before you deploy. After that, the territory is yours. This is not a support role behind an account team. You are the account team, the solutions architect, and the builder, compress
About Dot Collective We are a new generation consultancy based across UK and EU and founded on the premises of the engineering excellence and empowering people to make an impact. We work with all modern tech stacks and typically run agile scrum on all our projects. About you Are you passionate about data and its transformational powers? Do you like being able to make a huge difference in a limited period of time? We might be just the right place for you. Your key skills and capabilities: Engage with either AWS or GCP cloud ecosystems to ensure best practise development for new and existing solutions Build, deploy and manage Cloud Infrastructure through with IaC concepts Hands on experience with serverless services such as AWS’ S3, Glue or Lake Formation and GCP’s Cloud Functions, Big Query or Data Fusion Integrate native cloud services with 3 rd party solutions through the offered networking solutions Understanding of the Python ecosystem from local development to production environments Experience of DevOps approaches supported with Python Work within a delivery focused team using Agile methodologies Comfortable with Docker and some exposure to orchestration tools Review and implement security best practices within cloud environments We expect you to know how to architect, design, develop, deploy and operate a data platform and be a good leader for your team. Our promise to you We will always see you as a human being and will do our very best to support your needs and wellbeing – well-designed co-working and collaboration spaces, remote working patterns that work for you, parenting leave, sabbaticals and ability to work on personal projects. We believe that a geled team is worth its weight in gold – we will do everything we can to avoid breaking well-performing teams. Whilst continuity across every project is not always possible, we thoughtfully assemble high-performing, blended
Other cities to consider
More places hiring for this role
Get new production support sre analyst jobs in London, United Kingdom by email
Daily job updates · Unsubscribe anytime