We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity As a Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with a proven track record of building and scaling resilient infrastructure. What you'll do Platform Orchestration: Work on a large-scale K8s infrastructure platform, ensuring high availability and performance. Automation: Drive the evolution of internal tooling to streamline platform delivery. This role requires Experience: Solid hands-on background in DevOps, Site Reliability, or Infrastructure Engineering. Kubernetes Mastery: Deep internal knowledge of K8s primitives (Deployments, StatefulSets, Services) and hands-on experience writing custom Kubernetes Operators. Golang Proficiency: Proficiency in Go, specifically for infrastructure automation and systems programming. Operations-Heavy Mindset: A proven track record of managing production environments and handling high-severity incidents. Cloud Infrastructure: Hands-on experience with cloud-native scaling tools (e.g., Karpenter, Cluster API) and Day 1/Day 2 operations of K8s clusters. Tooling: Familiarity with Helm and GitOps workflows (e.g., ArgoCD or Flux). Please note that visa sponsorship is not available for this position. Fostering a diverse, welcoming and inclusive environment is important to us. We work hard to make everyone feel comfortable bringing their best, most authentic selves to work every day. We cele
Jobiba hiring network
Cloud Operations Engineer Jobs
2,329 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Role Redwood is scaling public cloud infrastructure and AI features across multiple product lines, and we need a FinOps Lead to bring rigor, visibility, and accountability to that spend. This is a senior individual-contributor role with the authority to drive cross-functional cost governance directly. This role owns the translation of raw cloud cost data and AI spent into the models, forecasts, and governance mechanisms that let engineering, product, and executive leadership make informed decisions. This is not a bill-monitoring role. You will build the cost attribution infrastructure that ties cloud/AI spend to specific product lines and, ultimately, to ROI involving architecture and engineering teams to make the right trade offs and decisions in line with the strategic roadmap. The role will report to the Senior Director, Product Engineering Operations, and work closely with Cloud Engineering, the CPO's org, and engineering leadership to create cost visibility and defensibility informing leadership on efficiency strategies. Your ability to be an effective communicator, collaborative team player, and analytical thinker will be keys to success in this role. Responsibilities Cost Visibility, Attribution & Optimization Own and continuously improve cost models that attribute AWS (and other public cloud) spend by product line, team, and environment Drive tagging governance and hygiene, define standards, audit compliance, and close attribution gaps that prevent accurate cost-per-product reporting Build shared cost allocation models to drive transparency into per-team cost drivers where there are prevalent savings plans, reserved instances and network charges Build toward feature-level cost attribution that connects infrastructure spend to product ROI, not just aggregate bill totals Create and manage optimization programs with achievable savings targets, including working across engineering and finance to rightsize, clean up and modernize infrastructur
PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. PagerDuty is seeking a Principal Solutions Consultant to join our talented, customer-focused team! You will be a key strategic leader and the "technical face" of PagerDuty across the US region. This is a high-impact role that sits at the intersection of Sales, Product, and Engineering. You won’t just be selling a platform; you will be partnering with CTOs, CIOs, and VPs of Engineering at the US’s most influential enterprises to redefine how they manage digital operations, resilience, and automation. You will act as a bridge between our customers’ long-term strategic needs and PagerDuty’s product roadmap. You will spend your time evangelizing the PagerDuty Operations Cloud, mentoring our high-performing technical sales teams, and acting as a trusted advisor to the C-suite on topics ranging from AIOps and SRE maturity to digital transformation and cloud migration. Key Responsibilities Act as a peer and advisor to customer CTOs and CIOs, helping them navigate complex digital transformations and operationalizing the PagerDuty Operations Cloud within their organizations. Represent PagerDuty as a thought lea
PagerDuty, Inc. (NYSE: PD) is the global leader in AI-first digital operations. By automatically detecting, diagnosing, and remediating issues, the PagerDuty Platform orchestrates AI agents and automated workflows with context from over 750 integrations. Trusted by approximately two-thirds of the Fortune 100 and nearly half of the Fortune 500, PagerDuty is the industry standard for organizations scaling resilient, autonomous operations. Notable customers include Chipotle, Cloudflare, Docusign, Fox, Nvidia, Salesforce, Spotify, Zoom and more. We are growing rapidly and hiring top talent with leading AI skills across engineering, sales, product, marketing, and beyond as we build the leading digital operations platform. About the Role PagerDuty is seeking a Principal Product Manager, Platform Security to own the strategy and execution of how we secure, harden, and defend our Operations Cloud platform. This role sits within our Product Development organization and reports to the Sr Director of Product, Platform & Partners. This is a senior individual contributor role. You'll bring the same rigor to security that a great PM brings to a product: deep customer empathy, structured threat modeling, clear risk tiering frameworks, and a bias toward measurable outcomes. You'll also be the connective tissue between Product, Engineering, IT, and Legal, ensuring security strategy translates into engineering execution and customer trust. This role owns the full lifecycle from recommendation to implementation to operations. You'll be the decision-maker on risk acceptance, control exceptions, and incident escalation in real-time. The ideal candidate has operated at the intersection of product management and security engineering in a later-stage B2B SaaS environment. You've owned security architecture decisions end-to-end, built security infrastructure, and have the credibility to influence both product roadmaps and engineering practices without formal authority. What You'l
A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Optimize cloud financial operations to maximize value from cloud investments, including rapidly growing artificial intelligence (AI) and machine learning workloads Provide actionable insights on cloud spend, SaaS license optimization, and emerging AI cost drivers, including model inference and usage-based consumption Implement tooling, tagging standards, and processes that improve cost visibility and optimization across cloud, SaaS, and AI workloads Monitor large language model API consumption and GPU-intensive infrastructure to identify cost trends, anomalies, and optimization opportunities Build financial models to forecast cloud, SaaS, and AI expenditures for budgeting cycles, commitment decisions, and vendor negotiations Design cost allocation, tagging, showback, and chargeback models that attribute spend to the teams, applications, and use cases driving it Educate engineering and business owners on cloud financial management practices th
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Technology, Data, and Intelligence (TDI) Team Okta is the leading independent identity provider. The Technology, Data, and Intelligence (TDI) organization is the engine that powers Okta's global workforce, providing the technology and systems that enable our employees to do their best work. The Staff Security Engineer Opportunity We are seeking a highly skilled and hands-on Staff Security Engineer with a DevSecOps & Issue Management focus to join the TDI Security team. In this role, you will be embedded directly within our technical environments , working side-by-side with engineering and operations teams to strengthen Okta’s security posture across infrastructure, cloud, and business systems. This is a tactical and strategic role; you will build a centralized Security Posture Analytics and Reporting capability that aggregates findings, remediation status, and risk data across Okta's security tools to improve visibility, accountability, and execution. The focus is on automating issue tracking and ownership assignment, identifying and routing unowned findings, partnering with remediation teams to accelerate closure, and developing remediation solutions where none currently exist. Longer term, leverage AI to simplify security posture management, increase trust in the data, automate analysis and prioritization, and provide actionable insights that help TDI proactively manage security risk at scale. The ideal candidate combines deep technical security e
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description: Okta - we are the World’s Identity Company. We don’t just protect logins; we secure the digital life of the Fortune 100. Built from the ground up in the cloud, Okta securely and simply connects people to their applications from any device, anywhere, at any time. Okta integrates with existing directories and identity systems, as well as thousands of on-premises, cloud and mobile applications, and runs on a secure, reliable and extensively audited cloud-based platform. Who are we looking for: We are looking for a Senior Software Engineer in Test. An individual who takes ownership and builds viable solutions. A team oriented individual who can demonstrate working independently, as an individual contributor, in a distributed working environment. Detail oriented and methodical in approaching tasks with excellent research and analytical skills. An ideal candidate will be someone who is passionate about automation & appling those automation skills in testing, cloud native, large-scale, mission-critical software in a fast-paced agile environment while partnering with cloud Infrastructure and operations teams. Okta engineering strongly believes in automated testing, and an iterative process to build high-quality next generation software. This role is mainly focused on automation of Cloud Infrastructure testing and supporting SREs, including bespoke solutions rolled out by Developer Productivity teams. The Quality Engineering team work
We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Platform Network Engineering Team Auth0 by Okta is an easy-to-implement authentication and authorization platform designed by developers for developers. We make access to applications safe, secure, and seamless for over 100 million daily logins worldwide. Our modern approach to identity enables this Tier 0 global service to deliver convenience, privacy, and security so customers can focus on innovation. The Senior Software Engineer Opportunity You will be part of the Platform Network engineering team responsible for all connectivity of Auth0. You will play a key engineering role as we evolve our network architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access. You will get to work with engineers throughout the engineering organization. What you’ll be doing Implement internal and edge networking infrastructure and design solutions that work at global scale and with multi-cloud and multi-region constraints. Carry cross-team initiatives from end to end: code reviews, design reviews, operational robustness, security hygiene, etc. Design and develop new services, tools, and automation to expose network functionality to other Okta engineering and operations teams. Research and implement solutions addressing cross-cutting concerns such as routing, failover, and scaling. Participate in the team’s on-call rotation. What you’ll bring to the role Have 3+ years of
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and proactive Security Engineer to help us build, maintain, and continuously improve the security posture of our rapidly growing ML infrastructure platform. As one of the first dedicated security hires at Baseten, you will work cross-functionally with engineering and operations teams to ensure we’re meeting the highest standards of confidentiality, integrity, and availability. You’ll have an opportunity to shape our security strategy and best practices from the ground up, influencing the way our platform handles sensitive data for both internal and external stakeholders. RESPONSIBILITIES Security architecture and design: Collaborate with engineering teams to design and implement secure systems and infrastructure, including cloud (AWS/GCP) environments and container orchestration platforms. Vulnerability management: Lead proactive vulnerability assessments, pen tests, and remediation efforts to ensure our products and infrastructure remain secure. Incident response: Develop and maintain incident response processes, including detection, analysis, containment, eradication, and post-incident reviews. Identity and access management (IAM): Oversee IAM strategies and tools to ensure the right people have the right level of access to our systems and data. Security compliance and audits: Work closely with operations to ensure compliance with relevant standards (e.g., SOC 2, ISO 27001) and
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. Role Description As the Director of Corporate Infrastructure, you will drive efforts to oversee the design, implementation, and operation of our corporate networks. This includes the IT Infrastructure DevOps, Security and SRE teams. You have a deep understanding of infrastructure as code, configuration as code & networking technologies, strong business acumen, lead through data and metrics, and operate with a high level of rigor and accountability. You are responsible for the reliability, scalability, sustainability, and efficiency of SoFi’s network infrastructure. As a member of the Corporate Infrastructure leadership team, you are directly accountable for the teams that are responsible for driving efforts to evolve and build our next-generation corporate network (to include infrastructure and wifi). This includes managing all aspects of our network - engineering and operations to improve user experience and performance, and also supporting the multi-terabit backbone network that interconnects edge PoPs, corporate offices, data centers, and cloud gateways. You will lead a diverse team through highly technical problems to achieve our overall strategic, operational, and financial goals. You will build and lead a high-performance team, displaying technical proficiency to support and scale the
We take play seriously. We’re looking for curious adventurers ready to find their party, fueled by imagination and drive to build what’s never been built before. At Hasbro and Wizards of the Coast, you’ll collaborate with passionate teams to reimagine our iconic brands and create experiences that spark joy, connection, and community through the magic of play. This is your chance to shape legendary play that lasts a lifetime. At Wizards of the Coast, we connect people around the world through play and imagination. From our genre-defining games like Magic: The Gathering® and Dungeons & Dragons® to our growing multiverse, we continue to innovate and build new ways to foster friendship and connection. That’s where you come in! The cloud platform underpins how services are built, deployed, secured, and operated across the organization. As a Principal Cloud Infrastructure Engineer, you will define and drive the governance model and automation strategy for cloud infrastructure. This includes setting policy-as-code standards, automating compliance and cost controls, and ensuring infrastructure is provisioned and operated through consistent, auditable, self-service pathways rather than tailored or manual processes. You will operate at both a strategic and hands-on level and will set direction for how cloud resources are governed and automated while also building the tooling and guardrails that make that direction real. Success in this role means engineering teams can self-service the build and operation of infrastructure with confidence that it is secure, cost-aware, and consistent by default, without slowing delivery down. What you'll do Cloud Governance Define and enforce policy-as-code standards (tagging, naming, encryption, network segmentation, access boundaries) across cloud accounts/subscriptions Establish guardrails using tools such as OPA, Sentinel, AWS Config/Service Control Policies to report on and prevent drift from
Artefact is a new generation of data service providers, specialising in data consulting and data-driven digital marketing. It is dedicated to transforming data into business impact across the entire value chain of organisations. We are proud to say we’re enjoying skyrocketing growth.The backbone of our consulting missions, today our Data consulting team has more than 400 consultants covering all Artefact's offers (and more): data marketing, data governance, strategy consulting, product owner… About the role ... As a Cloud Architect Manager at Artefact, you will be responsible for designing, implementing, and overseeing robust, scalable, and secure cloud solutions that support our clients’ data and technology ecosystems. You will collaborate closely with technical teams and business stakeholders to ensure cloud architecture is aligned with business needs, optimizing tool selection, cost efficiency, and scalability. Additionally, you will lead technical teams, mentor engineers, and contribute to Artefact’s growth by driving innovation and promoting best practices in cloud technology. What you'll be doing ... Lead the design and implementation of cloud infrastructure architectures (cloud, on-premise, hybrid) aligned with business requirements. Define and enforce best practices for cloud platform deployment and configuration, security, networking, operations and tools selection. Oversee the integration of cloud platforms with data pipelines, storage, and compute services. Standardize coding, deployment, and monitoring practices across cloud environments. Design and implement automated solutions to set up, monitor, and adjust cloud infrastructure efficiently. Lead problem-solving initiatives for complex technical challenges related to cloud environments. Provide technical leadership across multiple cloud projects, ensuring efficient resource management. Foster teamwork by mentoring, coaching and guiding cloud engineers and ensuring knowledge transfer. Manage client relat
Artefact is a new generation of data service providers, specialising in data consulting and data-driven digital marketing. It is dedicated to transforming data into business impact across the entire value chain of organisations. We are proud to say we’re enjoying skyrocketing growth.The backbone of our consulting missions, today our Data consulting team has more than 400 consultants covering all Artefact's offers (and more): data marketing, data governance, strategy consulting, product owner… About the role ... As a Cloud Architect Manager at Artefact, you will be responsible for designing, implementing, and overseeing robust, scalable, and secure cloud solutions that support our clients’ data and technology ecosystems. You will collaborate closely with technical teams and business stakeholders to ensure cloud architecture is aligned with business needs, optimizing tool selection, cost efficiency, and scalability. Additionally, you will lead technical teams, mentor engineers, and contribute to Artefact’s growth by driving innovation and promoting best practices in cloud technology. What you'll be doing ... Lead the design and implementation of cloud infrastructure architectures (cloud, on-premise, hybrid) aligned with business requirements. Define and enforce best practices for cloud platform deployment and configuration, security, networking, operations and tools selection. Oversee the integration of cloud platforms with data pipelines, storage, and compute services. Standardize coding, deployment, and monitoring practices across cloud environments. Design and implement automated solutions to set up, monitor, and adjust cloud infrastructure efficiently. Lead problem-solving initiatives for complex technical challenges related to cloud environments. Provide technical leadership across multiple cloud projects, ensuring efficient resource management. Foster teamwork by mentoring, coaching and guiding cloud engineers and ensuring knowledge transfer. Manage client relat
Get new cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime