We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe
Jobiba hiring network
Cloud Operations Engineer Jobs
2,288 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. Position: Software Engineer | Java Developer 5 Work Days Per Week Office Near Tai Seng MRT, Singapore Medical Benefits Provided Entitled to Yearly Bonus & Performance Bonus Job Requirements Bachelor's or Master's Degree in Computer Science, Software Engineering, or a related technical discipline Minimum 1 year experience of software development experience (Java Programming) is preferable for this post Strong understanding of computer science fundamentals (operating systems, network principles, data structures, and algorithms) Programming Languages: Spring Core, Spring Boot, Spring Security Cloud Platforms: AWS and Azure Frameworks: Common open-source frameworks and tools such as Kafka, RocketMQ, Dubbo, Zookeeper, and Redis. Deep knowledge of MySQL, including schema design, SQL optimization, and database scaling strategies. Job Responsibilities ( Software Developer / Java Backend Engineer) Design & Development Take ownership of the design, development, refactoring, and performance optimization of core system components, delivering high-quality and maintainable code. Technical Innovation & Problem-Solving Research, design, and implement innovative solutions to address complex business and technical problems, particularly in high-concurrency scenarios. System Architecture Contribute to system architecture decisions, focusing on scalability, high availability, and fault tolerance. Full-Lifecycle Part
About THG Ingenuity THG Ingenuity is a fully integrated digital commerce ecosystem, designed to power brands without limits. Our global end-to-end tech platform is comprised of three products: THG Commerce, THG Studios, THG Fulfilment. Each represents a single, unified solution, overcoming challenges and taking brands direct-to-consumer. Our client portfolio includes globally recognised brands such as Coca-Cola, Nestle, Elemis, Homebase, and Proctor & Gamble. Database Engineer (DBA) Company: THG Ingenuity Location: Manchester (Head Office) Reports to: Database Platform Manager Role Overview We are looking for a Database Engineer to help run and improve our database platform. Our estate runs on Google Cloud Platform following a recent migration to self-managed services. There is plenty still to do in the next phase of modernising the estate, and you will be hands-on in that work as well as in the day-to-day running of the platform. You will work alongside a small team of database engineers and closely with engineering and infrastructure teams, keeping our database environments secure, reliable and performant. There is real scope to grow here, and you will be supported to deepen your technical expertise and take on more as you do. The role participates in a rotating on-call rota and occasionally requires out-of-hours work to support deployments, maintenance or incident response. Key Responsibilities Database administration Install, configure, maintain and upgrade PostgreSQL and SQL Server environments across development, test and production Ensure database servers are securely configured, patched and compliant with operational standards Own backup, restore and maintenance strategies, and test recovery procedures regularly Maintain database security, access control and auditing practices Cloud and infrastructure &n
Role: Manager, Software Engineering Location: Hyderabad, India (Hybrid) Department: Product Development Reports to: Director of Engineering About GHX: GHX (Global Healthcare Exchange) is a leading healthcare technology company on a mission to simplify the business of healthcare and improve patient outcomes. Founded in 2000, GHX has built the GHX Global Network — the world’s largest cloud-based supply chain community connecting healthcare providers, suppliers, distributors, and partners to automate key processes, reduce costs, and increase operational efficiency. Its solutions span electronic trading, procurement automation, inventory and contract management, business intelligence, and data synchronization, helping healthcare organizations improve productivity and focus more on patient care. Over the years, GHX has enabled significant cost savings for the industry and continues to innovate with intelligent automation and AI-driven capabilities. Website: https://www.ghx.com/ LinkedIn: https://www.linkedin.com/company/ghx/ About the Role GHX is seeking a hands-on Engineering Manager with expertise in software engineering, release management, artificial intelligence and cross-functional collaboration. The work blends modern software practices with pragmatic AI to reduce manual exceptions, improve matching accuracy, and enhance the invoice processing end-to-end. Responsibilities Engineering Leadership Provide technical and people leadership for an engineering team. Manage, mentor, and grow a high-performing team of software engineers, fostering a culture of collaboration, innovation, and accountability. Partner with other engineering and cross‑functional leads to align priorities and ensure cohesive execution. Coordinate incident response for the team, including coverage and technical response to issue. Hands‑on Technical Leader Contribute to high‑impact coding when it adds meaningful value (foundational components, prototypes, critica
DeepIntent is the leading healthcare marketing platform, purpose-built to help marketers plan, activate, and optimize data-driven campaigns with speed and precision. Trusted by the world’s top healthcare brands and their agencies, DeepIntent uniquely unites media, identity, and real-world clinical data to power privacy-safe, omnichannel marketing across every screen. Backed by patented technology and proven outcomes, DeepIntent’s platform delivers measurable audience quality and script lift at scale. Learn more at www.deepintent.com . What You’ll Do: We are looking for a Software Engineer to help build and scale our core backend systems and data infrastructure. In this role, you will work hands-on to develop the foundational data pipelines, storage solutions, and robust architectures that drive our healthcare advertising solutions and support our core products, reporting APIs, and analytics initiatives. This is an excellent opportunity for a growth-oriented engineer to work with massive datasets, modern cloud technologies, and cross-functional teams to deliver high-performance, fault-tolerant solutions. Build & Operate: Develop, test, and maintain highly reliable, scalable, and cost-optimized distributed systems and data architectures. Enable Self-Service Data: Create automated ingestion, storage, and transformation pipelines that make it simple for downstream users to access and utilize new datasets. Empower Machine Learning: Design and operate data pipelines specifically tailored to support the complex workflows of our Data Scientists and Machine Learning Engineers. Drive Operational Excellence: Help implement and champion DataOps and DevOps practices across the team to ensure system reliability and smooth deployments. Contribute to Best Practices: Play an active role in establishing and refining formal data practices, architectures, and engineering standards for the organization. Cross-Functional Collaboration: Partner effectively with business stakeholders,
JOB TITLE SOFTWARE ENGINEER, TECHNOLOGY A CAREER WITH POINT72'S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology team is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU'LL DO We are looking for an experienced professional to work as part of the Finance Technology team. In addition to tactical development, you will be responsible for delivering and creating programs to modernize and scale the platform through technology upgrades, cloud technology adoption, and re-architecting business processes. You will work alongside world-class engineers, partnering directly with Finance stakeholders to understand existing workflows and deliver scalable replacements. This position offers deep domain exposure across FP&A, investor reporting, and compensation, and carries a high degree of ownership over the modernization roadmap. Specifically, you will: Build software applications and deliver software enhancements and projects supporting finance and investor processing. Work closely with business stakeholders to develop software solutions using test-driven and agile software development methodologies. Be responsible for system upgrades and features supporting resiliency and capacity improvements, automation and controls, and integration with internal and external vendors and services. Driving architecture of core platforms and accelerating modernizing leveraging AI tools Work with DevOps teams to manage and resolve operational issues and leverage CI/CD platforms while following DevOps practices within the team and projects. Con
About Eudia: Eudia is redefining the future of legal work with AI-powered Augmented Intelligence, enabling Fortune 500 legal teams to move faster, manage risk more effectively, and unlock new business value. Backed by $105M in Series A funding led by General Catalyst, we’re building a category-defining platform that blends AI-driven automation with human expertise, transforming legal from a cost center into a strategic growth driver. At Eudia, we move fast. Unlike traditional enterprise software, our teams ship solutions in days, not months—delivering real impact for some of the world’s largest companies, including Cargill, Coherent, DHL, and Duracell. We’re solving one of the most complex, unsolved challenges in AI: bringing trust, accuracy, and security to legal automation. We’re a team of builders, operators, and problem-solvers who are passionate about reshaping an industry that has long been resistant to change. If you’re looking for a place where you’ll be challenged, take ownership from day one, and work alongside some of the brightest minds in AI and legal —we’d love to meet you. About the Role: Are you interested in building a high-performance Agentic AI driven legal workflow system that supports our current and future scale of platforms? If so, we are looking for you to join our growing team in India. This person will work from our Bangalore office and actively collaborate with the Palo Alto team. We are looking for a Lead Software Engineer that will help develop the most secure, enterprise-grade software using and innovating the latest in Generative AI. The opportunity to tackle challenges in creating cloud-agnostic solutions, maintaining stringent security and compliance standards, and building scalable, resilient platforms for enterprise applications, data, AI, and search also exist while you will be able to routinely innovate on behalf of our customers, collaborating with the world's top
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Senior Engineering Manager , Twilio’s Segment team. About the job As a Senior Engineering Manager on the Twilio Segment Data platform/ pipelines team, you’ll build and scale systems that process several hundred thousands of data points per second. You will lead the development of high-scale ingestion and data processing systems You'll guide the team in designing, operating and maintaining complex distributed systems, ensuring reliability, performance, and cost-efficiency while querying petabytes of data for our customer data platform (CDP). Responsibilities In this role, you’ll: Own and deliver robust, high-scale routing experiences for Data platforms & pipelines for Twilio Segment. Champion team growth and success, prioritizing mentorship and individual development. Architect and operate always-available, complex distributed systems in cloud environments. Guide technical decisions, articulating trade-offs between cost, p
Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... GEDE (Global Edge and Domains Engineering) keeps GoDaddy's domains, edge, and aftermarket platforms running for millions of customers. Within GEDE, the Reliability Engineering team (Domains Production Engineering) is the group that gets the call when something breaks — and, more importantly, builds the systems that mean it breaks less often. We own the monitoring, compliance, and patching toolchain for the whole Domains infrastructure footprint, provide advanced incident support, and are actively modernizing how the org detects, diagnoses, and even auto-remediates issues with AI-assisted tooling! What you'll get to do... Build and evolve observability using Prometheus/Mimir, the Grafana LGTM stack, Elastic/OTEL, and Site24x7 — closing gaps across the org. Steer our cloud migration journey and bolster our efforts to keep the services reliable and performant. Own patching compliance and vulnerability remediation at scale across a mixed on-prem + AWS fleet, hitting hard SLA targets. Operate and extend our multi-tenant Kubernetes/ArgoCD platform, including the migration of core services. Contribute to our AI/automation initiatives: auto-generating runbooks from Prometheus alerts, ServiceNow change-risk scoring, and other tooling that reduces toil for the whole team Consult with partner dev teams on metrics, alert thresholds, and monitoring standards — this is a platform-enablement role, not just a ticket queue. Mentor other engineers on the team and help mature our operational practices. Your experience should include
Location Details: Remote, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings Join Our Team... GoDaddy’s Product Security team is looking for a Security Engineer who is passionate about helping build secure products and services used by millions of customers worldwide. In this role, you will work closely with engineering and platform teams to identify security risks, improve security practices, and help embed security throughout the software development lifecycle. You will have the opportunity to learn from experienced security professionals while contributing to initiatives that improve the security and resilience of our products, platforms, and cloud environments. The ideal candidate is curious, hands-on, eager to learn, takes ownership of their work, and excited to solve security challenges at scale! What you'll get to do... Identify, assess, and help remediate security risks across applications, infrastructure, cloud environments, and AI-enabled services Partner with engineering and platform teams to design, implement, and improve security controls throughout the software development lifecycle Conduct security reviews, threat modelling, and risk assessments to help build secure, resilient, and scalable systems Develop and enhance automation, tooling, and processes that strengthen security coverage, improve detection and response capabilities, and increase operational efficiency Investigate security findings, stay informed on emerging threats and technologies, and contribute to a culture of security awareness and continuous learning Your experience should include... 2+ yea
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The ability to monitor and assist our vehicles remotely plays a key role in our business strategy. As a Software Engineer, Video Streaming you will work on our in-house Teleoperations platform. You will work with a diverse team of engineers to build the core communication system as well as the cloud platform to connect vehicles and operators. This position involves broad technical understanding in networking algorithms, bandwidth estimation, rate control, computer networking, and real-time communication systems. The team is expected to deliver reliable solutions and license to 3rd party teleoperation usages. About the Work Design and implement an efficient pipeline with state-of-the-art video streaming techniques to deliver high priority real-time data stream Build an offline streaming simulation/emulation framework that can help to iterate the video streaming algorithm and predict online performance Test systems in real-w
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're searching for a highly motivated, technical leader to drive the engineering roadmap and innovation for our rack system software architecture. From firmware, kernel drivers, operating systems, networking, fabrics and associated user mode drivers + manageability software. You will work with component leads internally and engage with industry leading hyperscalar / cloud service providers on taking these products to market. What you’ll be doing: Drive the software end-to-end architecture for NVIDIA's rack-scale products Maintain deep understanding of the product portfolio and roadmap; translate forward-looking plans into clear, formal software requirements that anchor execution across the organization. Ensure high quality & reliable software; serving as a trusted architectural partner to teams requiring
We are looking for a Senior Forward Deployed Engineer to join the Customer Solutions team. You will be the technical authority embedded with our most complex customers, guiding them through deployment, architecture, onboarding, and the adoption of agentic development workflows. You bring deep hands-on experience from prior roles and use that depth to advise, design repeatable patterns, and drive customer outcomes end to end. You operate autonomously, own the technical success of your customers, and bring their experience back to shape how Coder builds and delivers. This is not an execution-only role. You are equally comfortable doing deep technical work with a customer and stepping back to design the repeatable pattern behind it. You are energized by ambiguity, motivated by customer outcomes, and capable of influencing organizational change alongside the technical work. This position is required to sit in the Eastern Time Zone. What You'll Do Serve as the primary technical authority for post-sales customers, guiding deployment architecture, environment design, and adoption of Coder across both human and AI development workflows Own onboarding engagements end to end, ensuring customers move from contract to productive adoption with speed and confidence Lead Get Well engagements where architecture decisions, rollout patterns, or organizational dynamics are limiting customer health or growth Help customers implement the technical and organizational changes required to adopt agentic development practices at scale Design and document repeatable delivery patterns across onboarding, architecture, and adoption that can scale across customer segments Design and recommend reference architectures tailored to each customer's cloud environment, security posture, and organizational constraints Translate customer environment complexity into clear guidance on networking, ingress, identity, and infrastructure patterns Anticipate technical and operational risks, escalate to the right
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: Join a team that builds robust, real-time distributed systems for a cutting-edge database. We care about performance, reliability, scalability, and most of all learning and having fun together. Whether you’re a seasoned coder or just getting started, if you’re passionate about technology and eager to learn, you’ll fit right in. Who we are: We show up to work, ready to collaborate and build technologies that make a difference, with people who genuinely care. We chase improvements such as tail latencies, bytes throughput, cache hit rate, and operational cost efficiency. We believe learning is ongoing and that even the most complex problems can have simple solutions. What You’ll Do: Collaborate with teammates to design and build database features that power AI applications. Learn how to tune performance and support reliability in distributed systems (don’t worry, we’ll guide you). Help Pinecone run smoothly on popular cloud providers. Take ownership of your work and grow your skills every day. Have fun. Who You Are: 5+ years of work experience - programming in Rust, Go, C++, or a comparable language. You’re genuinely curious about distributed systems and eager to dive deep into technical challenges. You approach problems with creativity and persistence, and you’re comfortable asking thoughtful questions or seeking feedback. You’re excited to learn, value constructive feedback, and appreciate mentorship. Bonus Points: You have hands-on experience with cloud platforms (AWS, GCP, Azure) or have demonstrated an ability to pick u
Get new cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime