About the Team Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We work on building robust, scalable, high performance components to support our distributed training workloads. Our priorities are to maximize the productivity of our researchers and our hardware, with the goal of accelerating progress towards AGI. Within Training Runtime, the Process Management team develops the distributed OS responsible for launching, coordinating, and supervising the large numbers of processes that make up modern training workloads. Our runtime sits beneath training frameworks and on top of research infrastructure, ensuring jobs run reliably across massive clusters while maintaining performance, stability, and observability. Success for us is measured by both system reliability and researcher velocity - enabling ideas to scale from experiments to production training runs. About the Role As a Training Runtime: Process Management Engineer , you will work on the software that ties thousands of computers together and exposes them as a unified system. This system has to serve individual researchers running multiple parallel experiments, as well as our largest training runs spanning 100’s of thousands and even millions of machines and accelerators. This requires easy to use, introspectable systems that can promote a fast debugging and development cycle, as well as relentless optimization for scale while maintaining stability and performance throughout. You will work primarily in Rust , building high-performance asynchronous systems with a strong emphasis on performance, correctness, and scalability. Working at this scale and at the frontier of AI development poses novel challenges. Out-of-the-box approaches often don’t work. The problems you will be working on are highly ambiguous and require strong design judgment as well as proficient execution to advance the state of our infrastructure. We’re loo
Jobs in United Kingdom
Computer Hardware And Networking in London
8 active opportunities · Updated September 2026
Showing
8 jobs
Explore current computer hardware and networking jobs in London. Filter by work mode, employment type, experience, department, date posted and distance.
As a Senior Sales Engineer , you will be the primary technical resource for our Account Executive team. You will share your product and technical expertise through presentations, product demonstrations, and technical evaluations (Proof Of Values). As the technical expert, you will work with clients to understand their requirements and pain points, then design the right solution for their business needs. During the sales cycle you will guide clients through trials and POVs, demonstrating Sumo Logic’s ability to meet and exceed their requirements and building a positive relationship that will provide continuous value to our customers. Finally, you will have the opportunity to work cross-functionally with our Product Management and Engineering teams to share your knowledge and experiences to ultimately improve our business and our customers’ success. We seek talent who wants to leverage their technical and people skills to help deliver solutions to clients directly and become a trusted advisor in the process. Above all else, you should be highly self-motivated and extremely curious to learn more about Sumo Logic and the vast problems that it can solve. Responsibilities Partner with the Account Executives to understand customer challenges and mains, and articulate Sumo Logic’s value proposition, vision, and strategy to customers Technically close complex opportunities through advanced competitive knowledge, technical skill, and credibility Understand and help orchestrate all phases of the sales cycle, including leading technical validations during the Proof of Value phase Be successful working with all levels of an organization, from executives down to individual developers and Site Reliability Engineers Deliver product and technical demonstrations of the Sumo Logic service Work cross functionally with Product Management and Engineering to improve the Sumo Logic service based on your experience with customers Requirements B.S. in Computer Science, Engineerin
About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform
About Dot Collective We are a new generation consultancy based across UK and EU and founded on the premises of the engineering excellence and empowering people to make an impact. We work with all modern tech stacks and typically run agile scrum on all our projects. About you Are you passionate about data and its transformational powers? Do you like being able to make a huge difference in a limited period of time? We might be just the right place for you. A little bit about the role Over the past 18 months, The Dot Collective has successfully established a dedicated Data Architecture practice , laying the foundation for robust, scalable, and innovative data solutions. This practice underpins our commitment to delivering high-quality, data-driven outcomes for clients across diverse sectors. As we continue to grow, we are seeking a Data Solution Architect to play a pivotal role in shaping and implementing cloud native data and AI platforms that make a real impact. Your Key Skills and Capabilities Work with technical teams to clarify and refine architecture so as to maintain consistency during deliver Recognise ways to describe the structure and behaviour of applications used in a business, with a focus on how they interact with each other and with business users or actors Help the customer articulate their requirements and formulate a suitable solution architecture Proven experience of designing and building data platforms. A strong understanding of modelling data and how to make data work at scale A grasp of modern data governance is also appreciated Have an understanding of Microservices and Serverless architectures. Be able to describe the structure and behaviour of the technology platform that underpins user applications and understand on-premise, Cloud hosting and end-user compute Knowledge of open source, open standards and cloud technologies An understanding of architectural concepts, methodologies and approa
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Please note that this is not a live role. We are grateful for your enthusiasm and interest in Cohere and encourage you to submit your resume for future consideration. Should a suitable position become available, we will review your application and reach out to discuss potential opportunities. Why this role? Design and implement novel research ideas, ship state of the art models to production, and maintain deep connections to academia. We have one of the highest ratio of compute to engineers in the world. We do not delineate strongly between engineering and research. Everyone will contribute to writing production code and conducting research depending on individual interest and organizational needs. We have all the compute, data, and talent available for you to do your best work. Please Note: We have offices in Toronto, London, Paris, San Francisco and New York but also embrace being remote-friendly! There are no restrictions on where you can be located for this role. As a Member of Technical Staff, you will: Design, build and scale AI systems for serving our users. Research, implement, and experiment with ideas on our supercompu
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Code-generating LLMs and autonomous agents are revolutionizing how software is built and tasks are automated. At Cohere, we’re pushing the boundaries of what’s possible with these technologies for enterprises, and we’re looking for a senior member for the Agent Code team. You’ll be at the forefront of research and engineering, driving the development of cutting-edge code LLMs and agent systems that can interact with the digital world to solve complex tasks with minimal human oversight. This role is hands-on and research-driven. You’ll dive into the latest literature on code LLMs and agents, experiment with frontier models, and collaborate with a team of talented engineers and researchers to build scalable, production-ready solutions. At Cohere, we blend engineering and research seamlessly—everyone contributes to both, depending on their interests and organizational needs. We provide access to world-class compute resources, data, and talent to ensure you can do your best work. Note: We have offices in London, Toronto, New York and San Francisco, but we’re also remote-friendly! This team operates primarily between E
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Advance the state of the art for model post training, ship state of the art models to production, and bridge the gap between research and production. We have one of the highest ratio of compute to engineers in the world. We do not delineate strongly between engineering and research. Everyone will contribute to writing production code and supporting our research effort depending on individual interest and organisational needs. We have all the compute, data, and talent available for you to do your best work. Please Note: We have offices in London, Paris, Toronto, San Francisco and New York but also embrace being remote-friendly! As a Member of Technical Staff, you will: Design and write high-performant and scalable software for training models. Consistently post-train the models to reach SOTA level performance. Coordinate with other specialist teams (Agentic, Code…) to produce models that have strong all encompassing performance. Craft and implement techniques to improve the performance and results of our training cycles both on the SFT and the RL regime. Research, implement, and experiment with ideas on our superco
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer on Detection & Response, you’ll help protect OpenAI’s most sensitive assets– including our intellectual property, customer data, and the infrastructure that supports them– by building and operating the systems we use to detect suspicious activity and respond effectively when it matters. You’ll work across endpoints, identity, cloud, hyperscale compute infrastructure, and datacenter-adjacent layers, partnering closely with security teams and infrastructure owners to define the telemetry and response requirements we need and building tooling and automation where it delivers the most leverage. In this role, you will: Build and evolve Detection & Response capabilities across OpenAI’s infrastructure, products, and research environments, with an emphasis on high-signal detection and reliable operational response. Engineer detection pipelines and tooling: develop rule lifecycle management, measurement/quality loops (coverage, precision, latency), tuning processes, and safe rollout patterns. Automate response and investigations by building workflows that reduce toil (triage, enrichment, containment, evidence capture) and improve time-to-understand/time-to-contain. Partner with other Security teams and system/infrastructure owners across the company to ensure new systems ship with the right telemetry, threat models, and response playbooks from day one. Define D&R requirements and drive visibility across endpoin
Get new computer hardware and networking jobs in London, United Kingdom by email
Daily job updates · Unsubscribe anytime