Jobs in United Kingdom

Production Support Sre Analyst in United Kingdom

96 active opportunities · Updated October 2026

Explore current production support sre analyst jobs across United Kingdom. Filter by work mode, employment type, experience, department, date posted and distance.

F
📍 London, England, United Kingdom· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment. Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done. Forward Networks is looking for a Systems (Sales) Engineer Do want to create a category and help build a special company? Do you want to help sell a platform that solves real networking problems? Join a company that has been in market 5+ years and has some of the top Federal agencies and F500/Global 2000 already buying and referenceable. If you have 5-10 years of wildly successful experience as a Sales Engineer selling to large enterprise accounts..you may be the one! We are building a special team and hope you consider us if you want to have the experience of changing the networking world as we know it. Responsibilities: Serve as the primary technical resource for the sales organization working with large enterprise accounts. Work with the sales team as the product advocate and key technical adviser Drive and manage the te

PythonGitGraphqlAI
O
📍 London, Greater London, United Kingdom· Full-time
✓ Quality checkedCompany trend -100%

About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineering teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in London. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Partner directly with enterprise customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteria. Design, build, and deplo

JavaScriptTypeScriptPythonJava
S
📍 United Kingdom· Full-time· Remote
✓ Quality checkedCompany trend -100%

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. You independently lead the most complex AI deployments Smartsheet undertakes. You own the full engagement lifecycle from technical discovery through production deployment through solutions org handoff. You architect multi-agent solutions, design client-personalized MCP resource packs, and build the Deployment Kits that transform how 200+ solutions consultants and partners operate. You mentor junior FDEs, drive the intelligence loop, and present field findings at weekly Applied AI strategy sessions. You are building a function, not filling a role. What You Will Do Lead complex, multi-system AI deployments end-to-end scope, architect, build, validate, and manage the customer relationship throughout. Own the AI workshop program for your pod, customize modules per customer, lead technical sessions, translate outputs into production requirements, evolve content from field learning. Architect multi-agent solutions selecting the right coordination pattern for each customer’s workflow characteristics and compliance requirements. Design client-specific and industry-specific MCP resource packs that serve personalized intelligence from the server so every connected AI surface gets smarter for that customer automatically. Own Deployment Kit quality for your pod. If a kit is not documented well enough for a solutions consultant with no engineering background to follow, it isn’t done. Lead Solutions Enablement Sprints: transfer AI deployment patterns to solutions consultants and partners with training materials and certification crite

JavaScriptTypeScriptPythonJava
S
📍 United Kingdom· Full-time· Remote
✓ Quality checkedCompany trend -100%

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. You independently lead the most complex AI deployments Smartsheet undertakes. You own the full engagement lifecycle from technical discovery through production deployment through solutions org handoff. You architect multi-agent solutions, design client-personalized MCP resource packs, and build the Deployment Kits that transform how 200+ solutions consultants and partners operate. You mentor junior FDEs, drive the intelligence loop, and present field findings at weekly Applied AI strategy sessions. You are building a function, not filling a role. What You Will Do Lead complex, multi-system AI deployments end-to-end scope, architect, build, validate, and manage the customer relationship throughout. Own the AI workshop program for your pod, customize modules per customer, lead technical sessions, translate outputs into production requirements, evolve content from field learning. Architect multi-agent solutions selecting the right coordination pattern for each customer’s workflow characteristics and compliance requirements. Design client-specific and industry-specific MCP resource packs that serve personalized intelligence from the server so every connected AI surface gets smarter for that customer automatically. Own Deployment Kit quality for your pod. If a kit is not documented well enough for a solutions consultant with no engineering background to follow, it isn’t done. Lead Solutions Enablement Sprints: transfer AI deployment patterns to solutions consultants and partners with training materials and certification crite

JavaScriptTypeScriptPythonJava
F
📍 London, England, United Kingdom
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! At Figma we believe design doesn’t end in a file or with your designer - it includes everything that goes into the product you ship: the production code that ties it all together, the systems and developer tools that make that code reliable, the context you provide to our AI agents, and the documentation that keeps everyone on the same page. The Roundtripping Area at Figma is redefining how Designers, PMs and Engineers collaborate. We are responsible for agentic workflows enabling ideation and prototyping on production codebases as well as accelerating the journey from design to code. In 2023, we launched Dev Mode, a suite of features that give developers everything they need to navigate design files and transform designs into code. In 2025, we introduced Figma’s MCP, accelerating how ideas get to production. Looking to the future, we aim to further reduce the barriers between design to code and code to design allowing ideation, prototyping and productionalising to happen seamlessly where it most makes sense. Our Roundtripping team is expanding, and we’re hiring AI Product Engineers across multiple levels in the UK. We’re looking for people who have built generative AI products and are eager to lead AI efforts end-to-end, from early ideas to production. Join us in shaping the future of AI at Figma. What you’ll do at Figma: Build and evolve Dev Mode our MCP tools and Make, Figma’s leading tools for dev/design collaboration Take part in building new 0→1 products within the agentic coding space Col

TypeScriptReactAI
S
📍 London, England, United Kingdom· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe

PythonSQLDockerKubernetes
S
📍 London, England, United Kingdom· Full-time
✓ High-confidence listing

£107K – £262K/yr

Quick readStrong listing-quality and freshness signals

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. So familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the SpaceXAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applicatio

SQLPostgreSQLMongoDBDocker
G
📍 Cambridge, United Kingdom· Full-time
✓ High-confidence listing

£73.5K – £99.5K/yr

Quick readStrong listing-quality and freshness signals

Senior: GBP 73,500 - 99,500 Staff: GBP 97,300 - 131,700 Subject to alignment to the responsibilities and duties of the role - we currently have multiple positions available at both Senior and Staff level. About the job Build the Linux distribution foundation that turns upstream software into trusted Graphcore platform releases. You will help create the Linux distribution that powers Graphcore AI systems. The team produces production-ready system images from proven upstream distributions. Your work will shape how releases are built, validated and prepared for deployment. You will strengthen the engineering path from upstream Linux software to dependable platform releases. You will build and improve automated pipelines, run established Linux test suites, and diagnose issues across build and validation flows. As the platform evolves, you will introduce controlled configuration and tuning changes with evidence-led validation. This is hands-on systems engineering with visible impact. You will help define reliable processes for a new team building a critical part of Graphcore’s platform. The team and culture You will join one of Graphcore’s newest engineering teams, helping shape its culture from the start. It is a small, co-located team where ownership matters and progress is visible. Work happens through close technical discussion, practical problem-solving and evidence-led decisions. Ideas are challenged openly, and the best path wins regardless of hierarchy. You will report to a leader who values technical credibility and invests in people’s growth. The team moves with pace, takes responsibility and changes direction when the evidence demands it. What we're looking for Strong practical experience working in Linux environments Experience building or maintaining automated CI/CD pipelines for reliable engineering workflows Proficiency in Python, Bash or similar

PythonCI/CDLinuxRest
P
📍 London, England, United Kingdom· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

ABOUT THE ROLE We are a leading streaming global fitness content company with studios around the world including London, revolutionizing the way people access and engage with fitness workouts. Our platform offers a wide range of interactive, live and on-demand fitness content that caters to users of all fitness levels, empowering them to stay fit and healthy from the comfort of their homes. As the Senior Manager of Broadcast Engineering, you will play a pivotal role in our mission to deliver high-quality, seamless, and engaging fitness content to our global audience. You will lead the Broadcast Engineering team based in London, ensuring the smooth operation and optimization of our broadcast infrastructure, content delivery systems, and broadcast equipment. This position reports to the Director of Global Production Technology. YOUR DAILY IMPACT AT PELOTON Oversee and guide the Broadcast Engineering team in designing, implementing, and maintaining an efficient and reliable broadcast studio facility to deliver the best member experience possible Collaborate with global broadcast engineering leads to maintain parity and system wide connectivity between facilities Manage the procurement, installation, and maintenance of all broadcast equipment, ensuring their proper functioning and readiness for live and on-demand fitness classes Collaborate with cross-functional teams, including Content Production Operations, IT, and Product, to streamline content workflows, improve efficiency, and enhance the overall broadcast transmission process Stay up-to-date with the latest trends, advancements, and emerging technologies in broadcast engineering and streaming to propose and implement cutting-edge solutions Lead the team in promptly addressing technical issues and incidents, minimizing downtime and disruptions to the streaming service Mentor and guide the Broadcast Engineering team members, fostering a culture of learning, growth, and innovation YO

RedisAWSRestAI
C
📍 United Kingdom· Full-time· Remote
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building integrations across multiple model providers. Experience with AW

TypeScriptReactAWSDocker
F
📍 London, England, United Kingdom
✓ Quality checked

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! At Figma, we're building tools that help teams design, build, and ship better products together. Figma Make enables teams to go from prompt to production code, and the Figma agent is a purpose-built design agent that edits files directly on the Figma canvas. We're looking for a deeply technical product leader to lead the four teams responsible for Figma's Code area: Code Context, Design to Code, Code to Design, and Code Platform. You'll manage and grow a team of Product Managers while partnering closely with engineering, design, data, and research leaders. This team builds the products and platform capabilities that help design and code move fluently together. You'll set the strategy for improving connecting customer codebases to Figma. This is a full-time role that can be held from our London hub or remotely within the UK. What you'll do at Figma: Lead the product aspects of the four teams driving Figma's Code area - Code Context, Design to Code and Code to Design, and Code Platform Set vision, strategy, and goals for how design and code stay in sync at Figma Own Design to Code and Code to Design workflows end to end, including quality measurement, improvement, and discoverability Guide the development of the code-context systems and translation primitives that power agentic workflows and round-tripping between design and code Manage and develop three Product Managers, hire the team's open Product Manager role, and establish a high bar for product management craft Partner with leaders across en

C
📍 United Kingdom· Full-time
✓ Quality checkedCompany trend -100%

As a Staff Software Engineer on Coder’s Agentic Engineering team, you’ll shape the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on while setting the team's technical direction. You’ll lead complex work, make sound architectural decisions, and help other engineers do their best work. What you’ll do here Set technical direction across Coder’s agent harness, integrations, and workflows. Design and build production systems in Go, with work across React and TypeScript where needed. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Lead complex projects from early ambiguity through production. Raise the engineering bar through design reviews, code reviews, and technical mentorship. Partner with Product and Design on clear, useful agent experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Deep experience building and operating production software systems. Strong hands-on experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. A track record of setting technical direction without formal authority. Strong architectural judgment and comfort working through ambiguity. Someone who makes the engineers around them better. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development environm

TypeScriptReactAWSDocker
PE
📍 United Kingdom· Full-time· Hybrid
✓ Quality checked

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments.

PE
📍 United Kingdom· Full-time· Hybrid
✓ Quality checked

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.

KubernetesAIGo
PE
📍 United Kingdom· Full-time· Hybrid
✓ Quality checked

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and implementation. On a daily basis, you will build end-to-end workflows, take them to production, and solve real world problems at the largest scale. You will have ample opportunity to contribute learnings from the field back to the Palantir AIP product suite. You will be on the forefront of extending Palantir's existing footprint and strategy into new markets and problem spaces opened up by Gen AI. Core Responsibilities Forward Deployed AI Engineers’ responsibilities look similar to those of a hands-on AI startup CTO: you’ll work in small teams to own delivery of high stakes projects with clients. A day’s work may include building LLM workflows on a large scale, interacting with customers to understand their needs and set their AI strategy, but the most impact will be driven by implementing solutions into the real world of our partner's organisations. Do you aspire to be an entrepreneur or an Applied AI leader? We believe Palantir is the best place — with the best colleagues — to learn how!

🔔

Get new production support sre analyst jobs in United Kingdom by email

Daily job updates · Unsubscribe anytime