About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineering teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in London. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Partner directly with enterprise customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteria. Design, build, and deplo
Jobs in United Kingdom
Reliability Engineer Iii in London
21 active opportunities · Updated October 2026
Showing
6 jobs
Explore current reliability engineer iii jobs in London. Filter by work mode, employment type, experience, department, date posted and distance.
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. The Vendor Monitoring Data team focuses on gathering external data and conducting risk analysis as part of Vanta's Vendor Risk Management (VRM) product. Our work provides comprehensive insights that help customers mitigate third-party risks effectively. As a Senior Fullstack Engineer, you'll drive complex projects across our technical stack while mentoring our talented engineering team. This role offers a unique opportunity to delve into a hyper-focused subject area: external attack surface scanning. You'll tackle unique technical challenges and contribute directly to Vanta's impact by helping customers continuously and comprehensively monitor risks across their vendor supply chain. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as a Senior Fullstack Engineer, Vendor risk management at Vanta: Identify, scope, and lead large technical projects, laying the groundwork for core products to evolve and scale into highly performant, reliable, and customizable systems Make effective tradeoffs that consider business priorities, user experience, and a sustainable technical foundation Engineer sophisticated monitoring and alerting systems to guarantee the reliability, speed, and integrity of our security data pipeline. Collaborate with security researchers to rapidly deploy new scanning techniques and
We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe
About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Role: Senior Technical Support Specialist (Enterprise Agentic AI) Company: Ema Unlimited Inc. Location: London Employment Type: Full-time/Remote 1. About Ema (Why Ema) Ema is building the world’s first Universal AI Employee — a production-grade agentic AI platform that automates real enterprise workflows across HR, IT, Finance, and Operations. Ema’s customers do not run demos. They replace mission-critical, manual business processes with agentic AI systems that operate across multiple SaaS tools, APIs, and human-in-the-loop workflows. In this world, support is not reactive . Support is production reliability, trust preservation, and system learning . At Ema, Senior Technical Support Specialists are operators of live AI systems , not ticket handlers. 2. Role Overview The Senior Support Engineer owns the health, reliability, and trustworthiness of Ema’s deployed agentic AI systems in production. This role sits at the intersection of: AI behavior Workflow orchestration Enterprise integrations Customer trust Engineering feedback loops This is: ❌ Not L1 / call-center support, ❌ Not a “just escalate to engineering” role, ❌ Not reactive firefighting only This is : A senior technical escal
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? As a Member of Technical Staff in Data Analysis and Evaluation, you will play a pivotal role in ensuring the quality, reliability, and performance of our large language models (LLMs). Your primary focus will be on designing and conducting data collection tasks, assessing and evaluating dataset quality, and analysing the robustness and generalisability of our models. You will work closely with cross-functional teams, including researchers, engineers, and data annotators, to conduct data-driven decision-making and improve the overall effectiveness of our AI systems. This role combines expertise in statistics, experimental design incl. human annotators, and machine learning to ensure that our models are trained on high-quality data and perform reliably across diverse scenarios. You will contribute to Cohere’s mission of advancing AI by ensuring our systems are robust, scalable, and impactful. Please Note: We have offices in London, Paris, Toronto, San Francisco, and New York, but we also embrace being remote-friendly! There are no restrictions on where you can be located for this role. As a Member of Technical Staff
Other cities to consider
More places hiring for this role
Get new reliability engineer iii jobs in London, United Kingdom by email
Daily job updates · Unsubscribe anytime