About Dot Collective We are a new generation consultancy based across UK and EU and founded on the premises of the engineering excellence and empowering people to make an impact. We work with all modern tech stacks and typically run agile scrum on all our projects. About you Are you passionate about data and its transformational powers? Do you like being able to make a huge difference in a limited period of time? We might be just the right place for you. Your key skills and capabilities: Engage with either AWS or GCP cloud ecosystems to ensure best practise development for new and existing solutions Build, deploy and manage Cloud Infrastructure through with IaC concepts Hands on experience with serverless services such as AWS’ S3, Glue or Lake Formation and GCP’s Cloud Functions, Big Query or Data Fusion Integrate native cloud services with 3 rd party solutions through the offered networking solutions Understanding of the Python ecosystem from local development to production environments Experience of DevOps approaches supported with Python Work within a delivery focused team using Agile methodologies Comfortable with Docker and some exposure to orchestration tools Review and implement security best practices within cloud environments We expect you to know how to architect, design, develop, deploy and operate a data platform and be a good leader for your team. Our promise to you We will always see you as a human being and will do our very best to support your needs and wellbeing – well-designed co-working and collaboration spaces, remote working patterns that work for you, parenting leave, sabbaticals and ability to work on personal projects. We believe that a geled team is worth its weight in gold – we will do everything we can to avoid breaking well-performing teams. Whilst continuity across every project is not always possible, we thoughtfully assemble high-performing, blended
Jobs in United Kingdom
Infrastructure Sourcing Operations Lead in United Kingdom
121 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure sourcing operations lead jobs across United Kingdom. Filter by work mode, employment type, experience, department, date posted and distance.
£107K – £262K/yr
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As an ideal candidate you have a good understanding of how highly scalable and reliable production infrastructure is built. Most of our backend infrastructure is written in Rust. So familiarity with a compiled language such as C++, Rust, or Go is highly beneficial. RESPONSIBILITIES: Build the SpaceXAI API that serves our models to developers worldwide Own the end-to-end system responsible for high-throughput inference, handling billions of tokens per minute with low latency and high availability, including model serving infrastructure, request routing, SDK development, rate limiting, observability, and efficient scaling BASIC QUALIFICATIONS: Expert knowledge of either Rust or C++ Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems Knowledge of service observability and reliability best practices Experience in operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDB PREFERRED SKILLS AND EXPERIENCE: Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM) Experience designing or building with agent SDKs and agent orchestration frameworks Experience with Docker, Kubernetes, and containerized applicatio
Why Sony Interactive Entertainment? Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work. Sony Interactive Entertainment (SIE) is the company behind the PlayStation brand. As a subsidiary of Sony Group Corporation, we’re part of a proud legacy of innovation and excellence. SIE is a dynamic technology company, delivering cutting-edge hardware and network services to more than 100 million people and an entertainment leader, home to some of the most beloved and recognizable intellectual properties (IP) in the world. Our role at SIE is to create and nurture the experiences under the PlayStation brand, a name synonymous with entertainment excellence and creativity. Join the team responsible for developing the tests, tools, automation, and infrastructure that help ensure the quality of the software toolchain used to build every PlayStation®5 game. You’ll collaborate closely with the ToolChain development teams to design tests and testing systems for new features as they are developed. You’ll also help design and implement internal tools, automation, diagnostics, and workflows that improve confidence in the quality, reliability, and maintainability of the product. This role sits within Quality Engineering. That means our work goes beyond running tests or finding bugs. We build software that helps teams understand risk, improve feedback loops, diagnose issues faster, and build quality into the development process as early as possible. Our toolchain is based on a private fork of the open-source LLVM project, continuously updated through an automated merge system. Thanks to our advanced continuous testing, we have a strong track record of identifying and reporting new bugs in the open-source project very quickly. This role offers a great opportunity to build a career in Software Engineering in Test and Quality Engineering, working on industry-leading development tools. You’ll receive mentorship from experts in the field and have
ABOUT THE ROLE We are a leading streaming global fitness content company with studios around the world including London, revolutionizing the way people access and engage with fitness workouts. Our platform offers a wide range of interactive, live and on-demand fitness content that caters to users of all fitness levels, empowering them to stay fit and healthy from the comfort of their homes. As the Senior Manager of Broadcast Engineering, you will play a pivotal role in our mission to deliver high-quality, seamless, and engaging fitness content to our global audience. You will lead the Broadcast Engineering team based in London, ensuring the smooth operation and optimization of our broadcast infrastructure, content delivery systems, and broadcast equipment. This position reports to the Director of Global Production Technology. YOUR DAILY IMPACT AT PELOTON Oversee and guide the Broadcast Engineering team in designing, implementing, and maintaining an efficient and reliable broadcast studio facility to deliver the best member experience possible Collaborate with global broadcast engineering leads to maintain parity and system wide connectivity between facilities Manage the procurement, installation, and maintenance of all broadcast equipment, ensuring their proper functioning and readiness for live and on-demand fitness classes Collaborate with cross-functional teams, including Content Production Operations, IT, and Product, to streamline content workflows, improve efficiency, and enhance the overall broadcast transmission process Stay up-to-date with the latest trends, advancements, and emerging technologies in broadcast engineering and streaming to propose and implement cutting-edge solutions Lead the team in promptly addressing technical issues and incidents, minimizing downtime and disruptions to the streaming service Mentor and guide the Broadcast Engineering team members, fostering a culture of learning, growth, and innovation YO
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team The Commerce Site Reliability Engineering team is responsible for the reliability, scalability, and day-to-day operation of the platforms that power GoDaddy's Commerce ecosystem. We build and operate shared infrastructure, support critical production systems, and partner closely with engineering teams to ensure services remain secure, resilient, and highly available. As a Senior Site Reliability Engineer, you'll join a team that values ownership, operational excellence, and continuous improvement. Engineers are empowered to identify problems, drive meaningful change, and influence how reliability is delivered across the broader Commerce organisation. From improving operational maturity and reducing toil to modernising delivery platforms and strengthening incident response practices, this team plays a key role in enabling engineering teams to move quickly and safely. You'll work closely with engineers across infrastructure, cloud, security, networking, and application teams while helping shape the future of reliability engineering at GoDaddy. This role offers significant opportunity to broaden your impact, develop technical leadership skills, and grow toward Staff and Principal engineering positions over time. What you'll get to do... Lead reliability and operational improvement initiatives across GoDaddy's Commerce platform, helping engineering teams build and operate services safely and at scale. Own critical production systems, drive incident response and post-incident improvements, and continuously raise the bar
We are seeking a highly technical and strategic Developer Relations Manager to join our team, with a focus on engaging developer ecosystems across emerging technology domains. In this pivotal role, you will work directly with software solution providers, developers, and industry professionals to foster the adoption of NVIDIA’s advanced AI and computing platforms. The ideal candidate brings a blend of deep technical expertise and commercial go-to-market experience, combined with a passion for developer advocacy and a talent for communicating how NVIDIA technology can solve complex, real-world challenges. NVIDIA is seeking a senior Developer Relations leader to accelerate the integration of NVIDIA technologies across the Energy & Utilities ecosystem in EMEA. This role will focus primarily on the electric power and utilities industry , working with leading software vendors, technology partners, utilities, engineering companies, and energy infrastructure providers to help integrate NVIDIA accelerated computing, AI, simulation, digital twin, and edge technologies into industry software platforms and solutions. The ideal candidate combines strong energy industry domain expertise , technical depth, partner engagement experience, and the ability to identify and develop strategic opportunities with Independent Software Vendors (ISVs). Approximately 80% of the role will focus on utilities, electrical power systems, grid software and related ISVs , with approximately 20% supporting Oil & Gas and adjacent energy applications . What You'll
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond to operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will design and develop systems that support the model lifecycle in production, including deployment orchestration, configuration management, operational automation, reliability, and capacity management. You will help transform complex operational processes into scalable platform capabilities that enable teams across OpenAI to deploy and manage models with greater confidence and less manual effort. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build and evolve the platform used to deploy, configure, and manage models across ChatGPT. Develop systems for deployment orchestration, model rollouts, operational visibility, and production readiness. Create abstractions and tooling that simplify complex infrastructure and improve the developer experience. Automate operational workflows, including incident detection, diagnosis, mitigation, and recovery. Improve the reliability, scalability, and efficiency of model deployments and the infrastructure that supports them. Build systems that support capacity planning, resource allocation, and infrastructure utilization. Partner with research, infrastructure, and product engineering teams to identify common chal
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Substrate is the team responsible for Palantir’s core production infrastructure — 100s of K8s clusters — from on-prem to the major cloud hyperscalers, whether they are internet-connected or air-gapped, small hardware footprint or large. As a Senior Software Engineer on Substrate, you will design and build Palantir’s managed Kubernetes product offerings across all these environments. You and your team will be responsible for bootstrapping and operating the entire fleet of K8s clusters with zero manual steps by building industry leading tooling and contributing to core CNCF components. You will also be responsible for ensuring scale, stability and security across a matrix of compliance regimes and hosting infrastructure types. Your team culture emphasizes engineering rigor and operational excellence at scale. This means issues in production should be pre-empted and deeply root-caused, and investments in automation and self-healing systems are key. If you’re excited about infrastructure at scale and working with Kubernetes, this is the right role for you.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed Enablement Engineers are embedded with centralised customer success teams to maximise the outcomes of our deployed products and workflows across all commercial customers from small-scale start-ups to large enterprises. This role balances high-level support responsibilities with the development of innovative tooling and infrastructure to scale customer enablement effectively. You are a resourceful, gritty and adaptable problem solver who is able to work both collaboratively and independently to resolve difficult and nebulous technical issues, as well as work productively with external customers to debug and resolve their problems. Palantir’s Customer Success team helps our customers build on Palantir’s Foundry & AIP Platforms to drive the workflows that power their most important business outcomes. In this role, you’ll leverage your problem-solving abilities, creativity, and technical skills to support and guide customer development teams, ensuring they can effectively build and optimise their workflows. You’ll have the opportunity to gain rare insight into and contribute to some of the world’s most important industries and institutions. Every day at Palantir is different: we’re constantly evolving to better respond to customer needs, and you will have the opportunity to contribute your creativity and problem-solving to internal processes and tools that define how we deliver business value to the customer with increasing efficacy and efficiency. Every Palantirian is encouraged to play to their personal strengths, so there is no “one size fits all” approach to engineering at Palantir. The scope of this role is intentionally broad so tha
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the Team: Bitstamp Exchange Platform Bitstamp made history in 2011 as the world’s first regulated crypto exchange. As a key part of the Robinhood family, the Exchange Platform team owns the full service lifecycle. We are the architects of a modern, high-velocity ecosystem that enables our global expansion, ensuring the world’s longest-running exchange remains unshakeable. The Role As a Senior Backend Engineer integrated into the Exchange Platform team, you will be a key driver of our modernization strategy. You will be an essential part of building the new generation infrastructure for low-latency services in the cloud, adopting cutting-edge technologies to drive our high-velocity ecosystem. This role is based in our London office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. Requires participation in an on-call rotation to support business needs. What You’ll Do Build Next-Gen Infrastructure: Architect and implement high-availability, low-latency cloud infrastructure that ensures performance and portability across our ecosystem. Modernize and Scale: Lead the transition of our services to modern, containerized environments, optimizing deployment and scaling workflows. Robinhood Ecosystem Integration: Manage and execute high-priority integration projects that align Bitstamp's backend with Robinhood's global infrastructure. Evolve the Stack: Identify and implement back
About the Role: The Program Manager executes Technology Capital Builds by creating tight alignment across Real Estate and Workplace Services, Corporate IT, Corporate Security and other partner teams to design and deliver technical solutions for capital build outs, including new office buildout and remodels, industrial labs, datacenters and secured facilities. This includes all low voltage, ISP connectivity, network infrastructure, audio visual, physical security and related IT scopes of work. In this role, you will: Ensure new sites launch with secure and reliable ISP connectivity, network infrastructure, and low voltage systems that are ready to support employees from day one. Deliver Capital Builds commitments through effective coordination across internal teams, construction partners, and vendors, achieving outcomes on scope, schedule, budget, and quality. Ensure AV systems across conference rooms, training spaces, all hands venues, digital signage, and wayfinding deliver a consistent and dependable user experience. Ensure access control, surveillance, and intrusion detection systems are integrated into the built environment and aligned with enterprise security requirements. Provide leadership with clear visibility into portfolio status, key decisions, dependencies, and emerging risks. Establish standards, drawing packages, specifications, and documentation that enable repeatable execution and operational consistency across the global portfolio. Ensure disciplined stewardship of procurement, budgets, and vendor investments across the Capital Builds portfolio. Identify and mitigate delivery risks early to protect project outcomes, business continuity, and operational readiness. Ensure all systems are commissioned, documented, and transitioned to support teams with clear ownership and support models in place. You might thrive in this role if you have: Strong project and program management capabilities. Strong knowledge of the architectural design process (schematic
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer on Detection & Response, you’ll help protect OpenAI’s most sensitive assets– including our intellectual property, customer data, and the infrastructure that supports them– by building and operating the systems we use to detect suspicious activity and respond effectively when it matters. You’ll work across endpoints, identity, cloud, hyperscale compute infrastructure, and datacenter-adjacent layers, partnering closely with security teams and infrastructure owners to define the telemetry and response requirements we need and building tooling and automation where it delivers the most leverage. In this role, you will: Build and evolve Detection & Response capabilities across OpenAI’s infrastructure, products, and research environments, with an emphasis on high-signal detection and reliable operational response. Engineer detection pipelines and tooling: develop rule lifecycle management, measurement/quality loops (coverage, precision, latency), tuning processes, and safe rollout patterns. Automate response and investigations by building workflows that reduce toil (triage, enrichment, containment, evidence capture) and improve time-to-understand/time-to-contain. Partner with other Security teams and system/infrastructure owners across the company to ensure new systems ship with the right telemetry, threat models, and response playbooks from day one. Define D&R requirements and drive visibility across endpoin
About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role This is where security meets innovation at enterprise scale. As a security engineer, applications at WRITER, you'll be building the security foundations that protect the AI systems powering some of the world's most recognizable brands. You'll work at the intersection of application security, AI infrastructure, and developer enablement—partnering with engineering teams to embed security into every line of code while ensuring our platform remains both powerful and trustworthy. The opportunity is massive: you'll help define how enterprise AI applications are secured, from threat modeling our LLM architectures to building automated security controls that scale across our growing platform. This isn't about saying "no"—it's about finding creative ways to say "yes, and here's how we do it securely." You'll tackle challenges that most security engineers never encounter: securing AI agents, protecting training data pipelines, and designing controls for systems that didn
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Join WRITER's security team as a staff detection and response engineer and help protect the AI infrastructure that's transforming how the world works. You'll build sophisticated detection systems that identify attacks targeting our AI platform, training data, and model deployments while creating automated response capabilities that scale with our explosive growth. This isn't just traditional security work – you're defending cutting-edge AI/AGI systems against adversaries who are evolving their tactics as fast as AI itself advances. This role combines hands-on security engineering with strategic thinking to stay ahead of novel threats that don't exist in textbooks yet. You'll be the operational arm of our security function, translating threat intelligence into real-time detections, coordinating incident response across multiple teams, and hunting for sophisticated attacks across GPU clusters and distributed training environments. If you're excited by the challen
Other cities to consider
More places hiring for this role
Get new infrastructure sourcing operations lead jobs in United Kingdom by email
Daily job updates · Unsubscribe anytime