Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Senior Security Operations Engineer you will: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Harden our cloud-native environments (AWS, OCI, GCP) by introducing secure by default designs and features into network, tooling, and processes Own and drive resolutions for enabling engineers to design, build, and use infrastructure securely at scale by deploying secure architectures using infrastructure-as-code and reusable code libraries Manage IAM / RBAC for cloud infrastructure, and partner with IT on streamling authentication/authorization to ensure unified access control across the board Deploy and operationalize some of the security services and tools (eg: SIEM, SOAR, domain monitoring, endpoint tooling, cloud security tooling) Respond to security incidents and harden environments post-incidents. Support control monitoring and remediation for compliance initiatives Gather and analyze security metrics to address security issues with cross-team dependencies Be a problem solver who is empathetic to developer concerns and will employ construc
Jobs in United Kingdom
Cloud Operations Engineer in United Kingdom
63 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud operations engineer jobs across United Kingdom. Filter by work mode, employment type, experience, department, date posted and distance.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Are you excited about building world class platforms that power mission critical software at a huge scale? Apollo enables autonomous management and continuous deployment of software, wherever it is. We’re taking SAAS to where SAAS has not gone before: from on-premise, to various cloud providers, to disconnected environments (air-gapped), to strict accreditation frameworks including IL-5 and FedRAMP, and to the edge. You can read more about the problem Apollo was built to solve on our blog or watch our Apollo demo day. As a Software Engineer on the Apollo platform you will build software at scale to transform how organisations around the world deploy software. You will be responsible for mission critical software powering deployment of software for both Palantir and its customers in the commercial and government space. You have a curious mind, high bar for engineering quality and ability to work in a dynamic environment where the best idea wins. Apollo Software Engineers are involved throughout the product lifecycle. From idea generation, to design and prototyping, to execution, and shipping. As a Software Engineer, you'll collaborate closely with technical and non-technical counterparts to understand our customers' problems and build products that solve them. Please note that this posting is open to all levels of experience.
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet's customers in Europe, the Middle East, and Africa expect our Customer Trust team to understand their compliance challenges, speak their language, and respond quickly to security assessments. We're looking for a Sr. Security Engineer I to lead Customer Trust operations across EMEA—responding to security questionnaires, managing vendor risk assessments, and building trusted relationships with enterprise customers in these regions. You will be responsible for questionnaire triage, completion, and queue management for EMEA customers. You'll work closely with EMEA sales teams, understand regional compliance requirements (GDPR, EU data protection, sector-specific frameworks), and ensure Smartsheet maintains a strong reputation for responsiveness and technical credibility in these high-value markets. You will work remotely from the UK and report to our Sr. Director, GRC Engineering, based in the US You Have 5+ years of experience in customer trust, vendor risk management, security assessment, or customer-facing security roles at SaaS or cloud platform companies. Proven experience completing and responding to customer security questionnaires, vendor assessments, and RFIs. Strong understanding of GDPR, EU data protection, and regional compliance requirements: Familiarity with data residency, data processing agreements, DPIAs, and how cloud services operate within EU regulatory frameworks. Knowledge of GRC frameworks: Working knowledge of SOC 2, ISO 27001, and compliance standards commonly referenced in EMEA asse
About the Team OpenAI’s Network Security team designs and operates the secure, reliable connectivity behind our offices, labs, campuses, cloud environments, people, and devices. We combine strong network fundamentals with automation, observability, and close partnership across IT, Security, Research, Applied, and business teams. About the Role As a Network Engineer, you will design, operate, troubleshoot, and automate secure, reliable networks across offices, labs, cloud connectivity, and production services. You will balance strategic platform work—architecture, standards, roadmaps, lifecycle planning, and automation—with responsive operations such as incidents, escalations, break/fix, and time-sensitive delivery. We’re looking for broad network engineers who meet users where they are, lead with curiosity, own outcomes end-to-end, move with urgency grounded in security, and iterate with purpose. You will turn operational signals and recurring reactive work into durable systems and standards. In this role, you will: End-to-end ownership of secure enterprise routing, switching, wireless, WAN, network services, and cloud connectivity. A deliberate balance of strategic platform improvement and responsive troubleshooting, change safety, incident response, and operational delivery. Purposeful iteration through software, APIs, Infrastructure-as-Code, Git workflows, testing, and CI/CD that reduces recurring reactive work. You might thrive in this role if you have: End-to-end ownership of secure enterprise routing, switching, wireless, WAN, network services, and cloud connectivity. A deliberate balance of strategic platform improvement and responsive troubleshooting, change safety, incident response, and operational delivery. Purposeful iteration through software, APIs, Infrastructure-as-Code, Git workflows, testing, and CI/CD that reduces recurring reactive work. Compensation, Benefits and Perks This is a position with OpenAI UK Ltd., which controls the hiring and manageme
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team The Commerce Site Reliability Engineering team is responsible for the reliability, scalability, and day-to-day operation of the platforms that power GoDaddy's Commerce ecosystem. We build and operate shared infrastructure, support critical production systems, and partner closely with engineering teams to ensure services remain secure, resilient, and highly available. As a Senior Site Reliability Engineer, you'll join a team that values ownership, operational excellence, and continuous improvement. Engineers are empowered to identify problems, drive meaningful change, and influence how reliability is delivered across the broader Commerce organisation. From improving operational maturity and reducing toil to modernising delivery platforms and strengthening incident response practices, this team plays a key role in enabling engineering teams to move quickly and safely. You'll work closely with engineers across infrastructure, cloud, security, networking, and application teams while helping shape the future of reliability engineering at GoDaddy. This role offers significant opportunity to broaden your impact, develop technical leadership skills, and grow toward Staff and Principal engineering positions over time. What you'll get to do... Lead reliability and operational improvement initiatives across GoDaddy's Commerce platform, helping engineering teams build and operate services safely and at scale. Own critical production systems, drive incident response and post-incident improvements, and continuously raise the bar
About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Senior Backend Engineer in GitLab’s Platform Enablement organization, you will help make GitLab easier to deploy, validate, and operate across environments. This team is centered on two major areas: Cloud Native deployment guidance and ephemeral environments. You’ll help shape GitLab’s next generation of self-managed deployment guidance as Reference Architectures evolve toward a model based on deployment patterns, workload characterization, component requirements, topology guidance, and scaling principles. You’ll also help improve production-like ephemeral environments so teams can validate changes e
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer on Detection & Response, you’ll help protect OpenAI’s most sensitive assets– including our intellectual property, customer data, and the infrastructure that supports them– by building and operating the systems we use to detect suspicious activity and respond effectively when it matters. You’ll work across endpoints, identity, cloud, hyperscale compute infrastructure, and datacenter-adjacent layers, partnering closely with security teams and infrastructure owners to define the telemetry and response requirements we need and building tooling and automation where it delivers the most leverage. In this role, you will: Build and evolve Detection & Response capabilities across OpenAI’s infrastructure, products, and research environments, with an emphasis on high-signal detection and reliable operational response. Engineer detection pipelines and tooling: develop rule lifecycle management, measurement/quality loops (coverage, precision, latency), tuning processes, and safe rollout patterns. Automate response and investigations by building workflows that reduce toil (triage, enrichment, containment, evidence capture) and improve time-to-understand/time-to-contain. Partner with other Security teams and system/infrastructure owners across the company to ensure new systems ship with the right telemetry, threat models, and response playbooks from day one. Define D&R requirements and drive visibility across endpoin
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building integrations across multiple model providers. Experience with AW
As a Staff Software Engineer on Coder’s Agentic Engineering team, you’ll shape the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on while setting the team's technical direction. You’ll lead complex work, make sound architectural decisions, and help other engineers do their best work. What you’ll do here Set technical direction across Coder’s agent harness, integrations, and workflows. Design and build production systems in Go, with work across React and TypeScript where needed. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Lead complex projects from early ambiguity through production. Raise the engineering bar through design reviews, code reviews, and technical mentorship. Partner with Product and Design on clear, useful agent experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Deep experience building and operating production software systems. Strong hands-on experience with Go. Experience with React and TypeScript. Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS. A track record of setting technical direction without formal authority. Strong architectural judgment and comfort working through ambiguity. Someone who makes the engineers around them better. Our tech stack Backend: Go, Postgres Frontend: TypeScript, React Infrastructure: AWS, Kubernetes Observability: Prometheus, Grafana CI/CD: GitHub Actions Bonus tacos if you have (Tacos? If you need an ice-breaker, ask how we say thanks by giving tacos!) Experience building coding agents, developer tools, or cloud development environm
We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe
About THG Ingenuity THG Ingenuity is a fully integrated digital commerce ecosystem, designed to power brands without limits. Our global end-to-end tech platform is comprised of three products: THG Commerce, THG Studios, THG Fulfilment. Each represents a single, unified solution, overcoming challenges and taking brands direct-to-consumer. Our client portfolio includes globally recognised brands such as Coca-Cola, Nestle, Elemis, Homebase, and Proctor & Gamble. Database Engineer (DBA) Company: THG Ingenuity Location: Manchester (Head Office) Reports to: Database Platform Manager Role Overview We are looking for a Database Engineer to help run and improve our database platform. Our estate runs on Google Cloud Platform following a recent migration to self-managed services. There is plenty still to do in the next phase of modernising the estate, and you will be hands-on in that work as well as in the day-to-day running of the platform. You will work alongside a small team of database engineers and closely with engineering and infrastructure teams, keeping our database environments secure, reliable and performant. There is real scope to grow here, and you will be supported to deepen your technical expertise and take on more as you do. The role participates in a rotating on-call rota and occasionally requires out-of-hours work to support deployments, maintenance or incident response. Key Responsibilities Database administration Install, configure, maintain and upgrade PostgreSQL and SQL Server environments across development, test and production Ensure database servers are securely configured, patched and compliant with operational standards Own backup, restore and maintenance strategies, and test recovery procedures regularly Maintain database security, access control and auditing practices Cloud and infrastructure &n
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity New Relic is looking for a Senior Business Value Engineer to join our Value Strategy team, specifically covering the EMEA region . In this role, you will be the primary value architect for our accounts across Europe, the Middle East, and Africa. You will bridge the gap between technical capabilities and business outcomes, helping our customers understand the financial and operational impact of the New Relic platform. You will act as a trusted advisor to both our internal sales teams and our customers' C-suite executives, driving the "Value Realization" process from initial discovery to long-term success. What you'll do Build Compelling Business Cases: Create and deliver Business Value Assessments (BVAs) including complex ROI modeling, TCO analysis, and value realization benchmarking. Strategize with Sales: Partner closely with Sales Leadership and Account Executives to develop deal strategies that emphasize business justification over feature-set comparisons. Quantify Technical Impact: Translate technical metrics (like MTTR, Error Rates, and Cloud Cost) into business KPIs (like Revenue at Risk, Subscriber Churn, and Operational Efficiency). Drive Value Realization: Ensure customers are co-creators in their value roadmap, moving beyond the initial sale to track and report on actual value achieved post-implementation. Educate & Enable: Mentor and train the broader GTM teams on value-based selling best practices to inc
Other cities to consider
More places hiring for this role
Get new cloud operations engineer jobs in United Kingdom by email
Daily job updates · Unsubscribe anytime