🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need. This is a hybrid position, based out of our New York City or London hubs. You'll report to our director of engineering. 🦸🏻♀️ What you'll do Technical
Jobs in United Kingdom
Infrastructure Sourcing Operations Lead in London
34 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure sourcing operations lead jobs in London. Filter by work mode, employment type, experience, department, date posted and distance.
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need. This is a hybrid position, based out of our New York City or London hubs. You'll report to our director of engineering. 🦸🏻♀️ What you'll do Technical
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Infrastructure Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for Palantir platforms, products, and deployments. You'll get to use your creativity to develop novel solutions to evolving challenges and automate processes wherever possible, using whichever tools are best for the job including industry-leading LLM and AI technology! As a Forward Deployed Infrastructure Engineer, every day is different! You will be developing software and providing high-quality support for software systems that are critical to solving our government’s greatest challenges. We strongly believe in engineering teams being responsible for the operations of their services in production. As such, you’ll work closely with forward deployed teams and product teams to participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
About the Team Our London-based team builds the backend systems that help ChatGPT scale reliably. We work on infrastructure close to the product, partnering with engineering teams to improve the performance, resilience, and operability of critical user-facing systems. Our work combines backend software engineering with distributed systems and production reliability. We build shared capabilities, improve high-traffic workflows, and make it easier to introduce new product functionality without compromising performance or availability. About the Role This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability. You’ll also own how those systems behave in production: how changes are rolled out, how issues are detected and diagnosed, and how recurring operational problems can be addressed through better software and system design. This is a strong fit for backend engineers who enjoy complex systems problems and want a direct connection between the infrastructure they build and the experience of ChatGPT users. In this role, you will: Design, build, and maintain backend systems supporting high-traffic ChatGPT experiences. Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely. Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow. Build and improve systems for asynchronous processing and other large-scale backend workloads. Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on-call, incident response, and root-cause analysis, and turn operational lea
About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and
£107K – £262K/yr
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Sandbox service team at SpaceXAI builds and maintains a secure, scalable system that gives our models safe, controlled access to computational environments. This infrastructure powers critical workloads across training and product, enabling models to run code, build software, interact with tools, and even control applications with user interfaces. We provision containers and virtual machines on large-scale clusters, granting models interactive control over these remote environments. Our work spans the full stack: from orchestrating massive jobs and resource scheduling at the cluster level, to fine-tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real-time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps. BASIC QUALIFICATIONS: Expert knowledge of Rust, C++ or Go Familiarity with Python Deep experience with either Linux or Windows systems (familiarity with both is a strong plus) Experience with virtualisation and containerisation technologies (e.g., cgroups, KVM, gVisor, QEMU) Solid knowledge of the networking stack COMPENSATION AND BENEFITS: £107,000 -
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Solutions Architecture team works with our largest, most complex users to understand their technical requirements and map those to Stripe technology. As a Solutions Architect, you'll partner with Sales to technically qualify new business opportunities, demonstrate the art of the possible with the Stripe Platform, and design robust technical solutions to enable payment transactions, manage money movements and simplify financial operational processes. What you’ll do You are an experienced technologist with a blend of technical depth and strong business consulting skills. You can write code, but prefer to spend your time working with users to create Stripe solutions to support customer business objectives in complex, mission-critical environments. You'll think strategically about the art of the possible for the user’s business, and be able to clearly articulate that in a way that informs and builds confidence in Stripe’s technology. You should be able to engage and motivate cross-functionally both internally and within our customer’s organisations at all levels, from Cx to Product Engineering. You'll have a track record of delivering exceptional customer results as part of a pre-sales team, but a history of deep technical work that gives you the ability to be credible with any audience. Responsibilities Engage with Chief Technical Officers, engineering leads, and other technical leads at key users to share technology roadmaps, demonstra
We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe
About Yondr Yondr is a disruptor. We challenge convention and simplify complexity. A global developer, owner operator and service provider of data centers, we deliver complex data center capacity needs for the world’s largest tech companies. Our exponential growth sees us looking for extraordinary people to help accelerate us towards our vision: a tomorrow without constraints. But we can’t do this without you. About the Role Yondr Group, a global developer, owner and operator of hyperscale data centres, is seeking an experienced lawyer to support its growing EMEA business. Reporting to the Legal Director EMEA, the successful candidate will act as the primary legal adviser to Yondr's design, construction and operations teams across the EMEA region. The role will support the full project lifecycle, from development, procurement and financing through construction, completion and operations. The successful candidate will work closely with project delivery, commercial, procurement and operational teams, supporting projects involving a range of delivery models, funding structures and stakeholders, including customers, lenders, investors, contractors, consultants, utility providers and suppliers. The role will be expected to operate independently on day-to-day matters, leading legal support across Yondr's EMEA construction and operations portfolio whilst managing legal risk and stakeholder requirements. This includes ensuring project documentation appropriately reflects contractual, leasing, funding and other stakeholder requirements. Minimum Qualifications Qualified solicitor with 6+ years' PQE with significant experience in primarily non-contentious construction, infrastructure, energy or data centre projects (preferred). Have trained at a reputable law firm or within the legal function of a reputable data centre, construction, infrastructure, energy or technology business. Experience negotiating and managing com
About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Threat Intelligence team protects OpenAI’s technology, people, research, and infrastructure by proactively identifying and disrupting adversaries who seek to compromise our systems or misuse our models. We investigate sophisticated threats, build tooling to scale and augment analysis, and deliver intelligence that shapes security strategy and equips leadership with timely, risk-aware insights. We combine technical depth, investigative rigor, and strong cross-functional partnerships to uncover threats and drive impact across OpenAI’s security and research organizations. About the Role As a Technical Threat Investigator at OpenAI, you will help protect the company from sophisticated adversaries targeting OpenAI and the broader ecosystem, as well as those attempting to misuse our models in support of cyber operations. This is a deeply investigative role. You will independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows. You’ll use these insights to proactively identify malicious activity and drive detection, disruption, enforcement, and safety improvements across the company. You’ll translate your investigative findings into durable solutions that scale impact. You’ll build and own lightweight tooling, automate where it matters, and create AI-assisted workflows to make investigations faster, more repeatable, and more effective over time. In this role, you will: Conduct deep, end-to-end investigations into sophisticated threat actors interacting with OpenAI’s models, products, and broader ecosystem. Think like an adversary — model attacker behavior, anticipate misuse patterns, and proactively hunt for, identify, and disrupt malicious activity. Leverage internal telemetry, OSINT, vendor data, a
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As Vanta continues its push into new markets, the EMEA Revenue Org is at the center of that growth — and this role is the dedicated enablement partner making that motion faster, sharper, and more scalable. As a Sr. Revenue Enablement Manager supporting our EMEA organization, you'll operate as a true strategic thought partner to Sales and Post-sales leadership, helping design and build the enablement infrastructure needed to win in complex sales environments. This is a high-autonomy role that shapes how Vanta's revenue teams learn, grow, and perform. This role will lead the execution of both global enablement initiatives and region-specific programs. You’ll work independently, build credibility with senior stakeholders, and help EMEA leaders reinforce new skills and behaviors with their teams. What you’ll do as a Senior Revenue Enablement Manager at Vanta: Partner closely with EMEA Revenue Leadership to understand business priorities, team performance, and enablement needs. Build trusted relationships with sales leadership and cross-functional partners across Revenue and the broader business. Diagnose gaps in industry, product, sales methodology, and role-specific skills across EMEA. Design and prioritize a regional enablement roadmap that connects global initiatives to the needs of EMEA teams. Support regional onboarding where needed. Create and deliver enablement programs, content, workshops, and other learning experiences that improve sales execution. Partner with EMEA leaders to reinforce enablement in the field and hold teams accountable for applying new knowledge and skills. Measure program effectiveness, gather feedback,
About Dot Collective We are a new generation consultancy based across UK and EU and founded on the premises of the engineering excellence and empowering people to make an impact. We work with all modern tech stacks and typically run agile scrum on all our projects. About you Are you passionate about data and its transformational powers? Do you like being able to make a huge difference in a limited period of time? We might be just the right place for you. Your key skills and capabilities: Engage with either AWS or GCP cloud ecosystems to ensure best practise development for new and existing solutions Build, deploy and manage Cloud Infrastructure through with IaC concepts Hands on experience with serverless services such as AWS’ S3, Glue or Lake Formation and GCP’s Cloud Functions, Big Query or Data Fusion Integrate native cloud services with 3 rd party solutions through the offered networking solutions Understanding of the Python ecosystem from local development to production environments Experience of DevOps approaches supported with Python Work within a delivery focused team using Agile methodologies Comfortable with Docker and some exposure to orchestration tools Review and implement security best practices within cloud environments We expect you to know how to architect, design, develop, deploy and operate a data platform and be a good leader for your team. Our promise to you We will always see you as a human being and will do our very best to support your needs and wellbeing – well-designed co-working and collaboration spaces, remote working patterns that work for you, parenting leave, sabbaticals and ability to work on personal projects. We believe that a geled team is worth its weight in gold – we will do everything we can to avoid breaking well-performing teams. Whilst continuity across every project is not always possible, we thoughtfully assemble high-performing, blended
£107K – £262K/yr
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. So familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the SpaceXAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applicatio
ABOUT THE ROLE We are a leading streaming global fitness content company with studios around the world including London, revolutionizing the way people access and engage with fitness workouts. Our platform offers a wide range of interactive, live and on-demand fitness content that caters to users of all fitness levels, empowering them to stay fit and healthy from the comfort of their homes. As the Senior Manager of Broadcast Engineering, you will play a pivotal role in our mission to deliver high-quality, seamless, and engaging fitness content to our global audience. You will lead the Broadcast Engineering team based in London, ensuring the smooth operation and optimization of our broadcast infrastructure, content delivery systems, and broadcast equipment. This position reports to the Director of Global Production Technology. YOUR DAILY IMPACT AT PELOTON Oversee and guide the Broadcast Engineering team in designing, implementing, and maintaining an efficient and reliable broadcast studio facility to deliver the best member experience possible Collaborate with global broadcast engineering leads to maintain parity and system wide connectivity between facilities Manage the procurement, installation, and maintenance of all broadcast equipment, ensuring their proper functioning and readiness for live and on-demand fitness classes Collaborate with cross-functional teams, including Content Production Operations, IT, and Product, to streamline content workflows, improve efficiency, and enhance the overall broadcast transmission process Stay up-to-date with the latest trends, advancements, and emerging technologies in broadcast engineering and streaming to propose and implement cutting-edge solutions Lead the team in promptly addressing technical issues and incidents, minimizing downtime and disruptions to the streaming service Mentor and guide the Broadcast Engineering team members, fostering a culture of learning, growth, and innovation YO
Other cities to consider
More places hiring for this role
Get new infrastructure sourcing operations lead jobs in London, United Kingdom by email
Daily job updates · Unsubscribe anytime