WHO ARE WE? We are a bunch of super enthusiastic, passionate, and highly driven people, working to achieve a common goal! We believe that work and the workplace should be joyful and always buzzing with energy! CloudSEK , one of India’s most trusted Cyber security product companies, is on a mission to build the world’s fastest and most reliable AI technology that identifies and resolves digital threats in real-time. The central proposition is leveraging Artificial Intelligence and Machine Learning to create a quick and reliable analysis and alert system that provides rapid detection across multiple internet sources, precise threat analysis, and prompt resolution with minimal human intervention. Founded in 2015, headquartered at Singapore, we are proud to say that we’ve grown at a frenetic pace and have been able to achieve some accolades along the way, including: CloudSEK’s Product Suite: CloudSEK XVigil constantly maps a customer’s digital assets, identifies threats and enriches them with cyber intelligence, and then provides workflows to manage and remediate all identified threats including takedown support. A powerful Attack Surface Monitoring tool that gives visibility and intelligence on customers’ attack surfaces. CloudSEK's BeVigil uses a combination of Mobile, Web, Network and Encryption Scanners to map and protect known and unknown assets. CloudSEK’s Contextual AI SVigil identifies software supply chain risks by monitoring Software, Cloud Services, and third-party dependencies. CloudSEK’s AIVigil is an AI-native Attack Surface Monitoring platform that continuously discovers, monitors, and secures exposed AI infrastructure, MCP servers, leaked AI credentials, vector databases, agentic workflows, and shadow AI across the internet. Key Milestones: 2016 : Launched our first product. 2018 : Secured Pre-series A funding. 2019 : Expanded operations to India, Southeast Asia, and the Americas. 2020 : Won the NASSCOM-DSCI Excellence Award for Security Product Company
Jobiba hiring network
Infrastructure Team Manager Jobs
4,815 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com . At Reddit, machine learning sits at the heart of how millions of people discover, connect, and engage with the world’s largest collection of human conversations. From powering personalized recommendations and search to optimizing advertising systems and marketplace dynamics, our ML engineers tackle some of the most interesting and impactful problems in large-scale applied machine learning. We hire Machine Learning Engineers across both our Consumer and Ads organizations, giving you the opportunity to work on a wide range of high-impact problems across the Reddit ecosystem. We are looking for Machine Learning Engineers who are excited to build systems end-to-end, from research and modeling to production deployment, — and who want to help shape the future of discovery, relevance, and monetization at Reddit. If you love working on complex, real-world ML problems at massive scale, this role is for you. What You’ll Work On As a Machine Learning Engineer at Reddit, you will design and build production ML systems that power core experiences across the platform, including: Personalized recommendations, search, and ranking systems that help users discover the most relevant content and communities Intelligent advertising systems including ranking, bidding, measurement, and optimization Content, Advertisers, and User understanding, from building foundational content/user representations to deriving insightful signals Large-scale machine learning pipelines, model serving infrastructure, and real-time decision systems Applied AI and
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: The Technology Risk & Controls Analyst operates as part of the second line Technology Risk & Controls function, working under the direction of the Technology Risk & Controls Manager, Tech Ops. The role supports the delivery of independent assurance over Technology Operations control environments, including infrastructure, cloud platforms, identity and access management, service management and core operational processes. The role is focused on executing high-quality control testing, audit support and remediation assurance activities in line with WPP standards and methodologies. The Analyst works closely with Technology Operations teams, Financial Risk & Control colleagues and auditors to evidence control performance, support audit readiness (including SOX 404) and contribute to a consistent, well-documented and sustainable control environment. What you'll be doing: Perform control design and operating effectiveness testing across Technology Operations processes, including change management, access management, incident/problem management, backup and resilience controls,
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: At WPP, technology is at the heart of everything we do, and it is the Technology Operations teams mission, as part of Enterprise Technology , to enable our stakeholders to collaborate, create and thrive. Enterprise Technology is undergoing a significant transformation to modernise ways of working, shift to cloud and micro-service-based architectures, drive automation, digitise colleague and client experiences and deliver insight from WPP’s petabytes of data. This role will carry out the effective and efficient everyday technology operations for WPP ET. A trusted pair of hands to deal with level 1 and 2 issues as they present to the IT Service Desk and a trusted resource for Infrastructure and Management personnel to assist with project work when needed. The role will report into the Enterprise Technology Operations Lead and work closely with other teams within Enterprise Technology. What you'll be doing: Deliver world class, on-site support services to WPP employees, agencies, and visiting clients, operating within predefined structur
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: The role is responsible for leading and overseeing end-to-end cloud operations, ensuring the availability, reliability, security, performance, and resilience of cloud platforms and services. It manages service monitoring, incident response, and incident resolution for production applications and cloud infrastructure, while ensuring operations teams are skilled and enabled to execute cloud-related requests with speed and diligence. The role also provides governance over a large third-party managed services organisation, ensuring delivery against agreed KPIs, SLAs, and operational commitments through effective service reviews, metrics reporting, performance management, and continuous service improvement. What you'll be doing: Responsible for overseeing cloud operations, driving operational excellence, and improving the operational landscape through automation and AI-driven solutions delivered by internal resources and third-party partnerships. Product: Work with product and engineering teams to define operational support patterns for each cloud product. Collaborate with business, archi
Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. Why is this a critical role at OneTrust? What is the challenge / type of challenges someone in this role will have the opportunity to take on? What impact does someone in this role have the opportunity to make within the company, for our customers and on a larger scale in the evolution of privacy and trust? An awareness of current issues affecting the industry and its technologies Create innovative, scalable, fault-tolerant software solutions for our clients and customer base Expand existing software to meet the changing needs of our key demographics What does this person do each day/each week? Describe a true to life day in the life for someone in this role. What do they do and how do they do it? Goal is to paint a real, genuine picture so candidates can see themselves in the role. Support production customers by monitoring and maintaining our cloud application & cloud infrastructure hosting it Build scripts for operational automation and incident response Handle processes surrounding cloud application deployment for our agile release Work with the monitoring, tuning, maintenanc
Who we are Graviton Research Capital is a privately funded quantitative trading firm. We trade across a multitude of asset classes and trading venues using a diverse range of concepts, from time series analysis and stochastic models to machine learning and statistical inference. We analyse terabytes of data to identify pricing anomalies and drive innovation in financial markets. Role Overview We are looking for a Program Manager who thrives at the intersection of rigorous engineering and predictable delivery. You will not just "manage tasks" — you will orchestrate the development lifecycle for mission-critical systems. Your goal is to ensure that our elite engineering teams can focus on high-performance code while you own the execution strategy, dependency mapping, and release discipline. Key Responsibilities Lead Agile ceremonies (Sprint Planning, Stand-ups, Retrospectives) tailored for deep-tech engineering teams. Transform high-level trading requirements into granular, executable backlogs. Own capacity planning and burn-down metrics to provide high-visibility delivery timelines. Navigate the complex interplay between engineering teams (e.g., Connectivity, Core Infrastructure, Simulation) to prevent bottlenecks. Build and maintain advanced Jira dashboards, automated roadmaps, and Confluence documentation that serve as the "single source of truth" for stakeholders. Proactively identify technical debt, architectural blockers, or resource gaps that threaten release stability. Continuously refine Agile methodologies to suit low-latency, performance-sensitive development cycles (where "Definition of Done" includes rigorous performance benchmarking). Eligibility & Required Skills 5+ years of experience as a TPM, Program Manager, or Scrum Lead in a product-engineering environment (HFT, FinTech, Networking, or Kernels/Systems). A strong grasp of the software development lifecycle for high-performance systems. While you won't write code, you must understand concepts li
Role Overview You’ll be the Principal Software Engineer driving the next generation of a large-scale enterprise SaaS platform. In this role, you combine deep hands-on engineering with high-impact technical leadership, shaping how cloud-native and AI-enabled products are designed and built. You’ll design and deliver secure, scalable, serverless systems on AWS using TypeScript and Node.js, modernize critical platform components, and set the technical direction for multiple teams. You’ll also lead how AI capabilities are integrated across the product ecosystem, ensuring they are transparent, observable, and compliant. If you enjoy system-level thinking, complex distributed architectures, and mentoring senior engineers while still staying close to the code, this role gives you company-wide impact and the opportunity to define the long-term technical vision. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and delivery of secure, scalable, serverless applications on AWS using TypeScript/Node.js. Define and evolve the platform architecture, driving modernization, performance, resilience, and maintainability. Design and operate distributed, event-driven systems using services like Lambda, DynamoDB, Aurora, S3, and EventBridge. Shape and implement AI-enabled solutions, embedding governance, observability, and responsible AI practices into the platform. Own Infrastructure as Code (e.g., Terraform, AWS CDK, CloudFormation) to reliably provision and manage cloud infrastructure. Mentor senior engineers, influence technical decisions across teams, and clearly communicate complex concepts to diverse stakeholders. These are the essentials you’ll need to get an interview Extensive experience (typically 12+ years) building secure, production-grade software systems. Proven track record architecting and delivering cloud-native, serverless applications on AWS. Strong expertise in Node.js, TypeScript, REST API design, and at leas
Here's a summary of the role: Do you love building scalable cloud platforms and solving complex engineering problems with modern technologies? As a Senior Software Engineer at Diligent, you'll design and deliver high-performing , serverless applications that power our global SaaS platform. You'll work extensively with TypeScript, Node.js, AWS, and event-driven microservices, owning services from design to deployment and production monitoring. This is an opportunity to influence technical decisions, mentor engineers, and explore how AI can transform software development and engineering productivity. If you're passionate about cloud-native architectures, distributed systems, and building software that scales to millions of users, we'd love to meet you. Here's a breakdown of what you'll do (not all of it, just the important stuff): Design and build scalable backend services and event-driven microservices using TypeScript and AWS. Develop secure APIs and integrations that power reporting, analytics, and dashboard experiences. Build and maintain serverless solutions using AWS services such as Lambda, EventBridge , SQS, and DynamoDB. Drive engineering excellence through testing, observability, automation, and production readiness practices. Contribute to infrastructure-as-code and CI/CD pipelines using AWS CDK and modern DevOps practices. Mentor engineers, participate in architecture discussions, and champion the use of AI tools to improve development efficiency. These are the essentials you'll need to get an interview: 6-8 years of professional software engineering experience. Strong experience with TypeScript, Node.js, and modern backend development patterns. Hands-on experience building cloud-native applications on AWS. Strong understanding of serverless architectures and event-driven microserv
About Scale AI At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. Scale Frontier Data is the organization behind the training and evaluation data that frontier labs depend on. We build the systems, tooling, and expert workflows that turn hard human expertise into signals that models can learn from, across reasoning, coding, agentic tool use, and domain expertise. Reinforcement learning environments are now the center of gravity for that work: the difference between a model that demos well and a model that reliably completes long-horizon work is almost always the quality of the environments and reward signals it was trained against. Responsibilities As a Staff Software Engineer, RL Environments, you'll own the technical foundation for how Scale builds, runs, verifies, and delivers RL environments at scale. An RL environment is a real piece of software: a containerized world with real dependencies, real state, real tools, and a grader that has to be correct even when the agent is creative about breaking it. Building one is a full-stack engineering problem. Building thousands of them reproducibly, cheaply, with trustworthy reward signals and throughput measured in millions of rollouts is a systems problem that very few people have solved. You'll work on both. You'll design the platform: sandboxed execution, environment packaging and versioning, rollout orchestration, trajectory capture, verifier frameworks, and the authoring surfaces that let engineers and domain experts produce environments without reinventing infrastructure each time. And you'll go deep on the environments themselves by instrumenting real applications, designing task suites that expose specific capability gaps, and building graders that
At Scale, our mission is to develop reliable AI systems for the world's most important decisions. For 10 years, Scale has provided the high-quality data and full-stack technologies that power the world's leading models, and has helped enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. Scale's internship is not a side project. Interns own real, shipped work on the same roadmaps as full-time engineers, with mentorship from world-class talent and a culture that values ownership, speed, and truth-seeking. Many of our interns return as full-time Scaliens. Example Projects Build reinforcement learning and post-training data pipelines that power frontier model development Develop evaluation infrastructure that measures model reliability for enterprise and public sector customers Ship agentic AI applications and the tooling that makes them observable, testable, and safe to deploy Ship tools that accelerate the growth of new qualified contributors on Scale's platform Build fraud-detection systems that remove bad actors and keep Scale's contributor base safe and trusted Use models to estimate the quality of tasks and contributors, and guarantee quality on requests at large scale Devise advanced matching algorithms that pair contributors to customers for optimal turnaround and accuracy Create optimized and efficient UI/UX tooling, in combination with ML algorithms, for 100k+ contributors completing billions of complex tasks Develop new AI infrastructure products to visualize, query, and explore Scale data Requirements A graduation date in Fall 2027 or Spring 2028 with a Bachelor's degree (or equivalent) in a relevant field (Computer Science, EECS, Computer Engineering, Statistics) Available for a Summer 2027 internship (May/June start dates) in San Franci
Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform providing APIs for knowledge retrieval, inference, evaluation, and more. We are seeking a strong Senior Full-Stack Engineer to help us build, scale, and refine our rapidly growing product. The ideal candidate is deeply grounded in software engineering best practices and experienced in developing and scaling modern web applications end-to-end. You will work across the stack—from React/TypeScript frontends to Python-based backends—while integrating with LLMs and machine learning systems. You will solve complex challenges in scalability, reliability, and product experience while owning significant product areas in a fast-paced environment. What You’ll Do Own major full-stack product areas , driving features from design through production deployment. Build modern frontend experiences using React and TypeScript, ensuring performance, usability, and responsiveness. Develop reliable backend services in Python, working with distributed systems, data pipelines, and ML/LLM components. Integrate with LLMs, vector databases, and AI infrastructure to power intelligent product experiences. Deliver experiments and new features quickly , maintaining high quality and tight feedback loops with customers. Collaborate across product, ML, and infrastructure teams to shape the direction of Scale GP. Adapt quickly —learning new technologies, frameworks, and tools as needed across the stack. Ideal Experience 5+ years of full-time engineering experience , post-graduation. Strong experience developing full-stack applications using React, TypeScript, and Python . Experience scaling or shipping products at high-growth startups . Familiarity with LLMs, vector databases, embeddings, or other modern AI tooling (tinkering or production experience welcome). Proficiency with SQL and modern API development. Experience with Kubernetes , containerization, and microservice architectures. Experience working with at leas
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Software Engineer, Data, you will design, build, and operate the next generation of our data platform and products – going beyond ID to power a networked digital identity – while keeping member privacy, security, and reliability at the core. What you’ll do: Build and operate scalable, reliable data systems and pipelines – from ingestion to modeling to visualization – so Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner. Develop and maintain end-to-end data products and pipelines (batch and/or streaming) that collect, clean, transform, and model data, and own the infrastructure that powers them to unlock new business use cases and reporting. Implement and maintain infrastructure-as-code, CI/CD, and shared developer tooling for data products (e.g., Pulumi/Terraform, GitHub, orchestration tools like Dagster/Airflow) to make it easy and safe for teams to build, test, and ship changes across environments. Improve the security, compliance, and cost posture of the data stack through robust dependency management, IAM and secrets hardening, observability, and performance/cost optimizations. Partner with product and other stakeholders to uncover requirements, make architectural decisions, and continuously improve our data platform and processes. How you’ll measure success: Data reliability & SLAs: % successful pipeline runs, adherence to freshness SLAs for core datasets, and reduction in data-related incidents impacting stakeholders. Platform quality & efficiency: Reductio
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for a strategic and execution-oriented Senior Director, Revenue Operations to lead and scale the operational backbone of our B2B organization. Sitting within B2B Operations, this role will own the strategy, architecture, and optimization of our revenue systems, processes, and analytics across Sales, Customer Success, and Marketing. You will serve as a key partner to B2B leadership, driving operational rigor, scalable infrastructure, and data-driven decision-making to accelerate revenue growth. You will define the roadmap for revenue systems, lead cross-functional initiatives, and build the foundation for long-term scale. What you'll do: Define and execute the B2B Revenue Operations roadmap in alignment with C1 growth objectives Act as a strategic partner to B2B leadership on forecasting, pipeline health, performance metrics, and operational investments. Establish scalable processes that improve conversion, velocity, forecasting accuracy, and revenue predictability at scale. Lead the redesign of Salesforce to support complex B2B sales motions, with hands-on responsibility for system configuration, reports, and dashboards Architect and optimize the full revenue tech stack (Salesforce, Outreach, ZoomInfo, HubSpot, Gong, LinkedIn Sales Navigator, etc.) Maintain data quality (deduplication, enrichment, normalization), build and evolve reporting frameworks, and troubleshoot integration issues across the revenue tech stack when they arise. Create and maintain internal documentation, runbooks, and training materials; support enabl
TEGNA Inc. helps people thrive in their local communities by providing the trusted local news and services that matter most. With 64 television stations in 51 U.S. markets, TEGNA reaches more than 100 million people monthly across web, mobile apps, streaming, and linear television, while also maintaining a strong global presence in India with offices in Bangalore and Chennai that support technology, product, and business operations initiatives. Together, we are building a sustainable future for local news. Senior DevOps Engineer About TEGNA TEGNA Inc. (NYSE: TGNA) helps people thrive in their local communities by providing trusted local news and services. With 64 television stations across 51 U.S. markets, TEGNA reaches more than 100 million people monthly across digital, mobile, streaming, and television platforms. We are focused on innovation, technology excellence, and building scalable solutions that create meaningful impact. Position Overview TEGNA is looking for a highly skilled Senior DevOps Engineer with strong expertise in AWS, Kubernetes, and Infrastructure as Code to design, automate, and manage scalable cloud infrastructure. The ideal candidate will have hands-on experience operating Kubernetes workloads in production, building CI/CD pipelines, and implementing monitoring and security best practices. This role requires deep technical expertise, strong troubleshooting skills, and the ability to work in fast-paced, distributed environments. You will play a key role in ensuring platform reliability, automation maturity, and production stability across cloud-native microservices systems. What You’ll Do Design and manage cloud infrastructure using Infrastructure as Code (AWS CDK, CloudFormation, Terraform). Build and maintain CI/CD pipelines using GitHub Actions and Jenkins to enable automated and reliable deployments. Deploy, manage, and scale Kubernetes clust
Get new infrastructure team manager jobs by email
Daily job updates · Unsubscribe anytime