Jobiba hiring network

Cloud Operations Engineer Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
1mo ago

About the Team Security is at the foundation of OpenAI's mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect model weights, customer data, and critical systems across multiple cloud environments. The team partners across OpenAI, including Applied Engineering, Research, IT, Security, Infrastructure, and Engineering, to provide secure and scalable platforms for identity, access management, permissioning, orchestration, and safe AI research. About the Role We’re looking for an engineering leader to lead Identity Infrastructure Engineering, the team building the systems that govern and scale access across OpenAI’s research, engineering, and internal platforms. This role sits at the center of cloud infrastructure, identity, software engineering, and security-critical operations. You’ll lead engineers building control planes, policy systems, workload and agent authorization patterns, infrastructure-as-code, and operational foundations that help OpenAI move quickly while keeping access reliable, auditable, least-privileged, and safe under failure. The ideal candidate has led teams responsible for large-scale, mission-critical infrastructure. They can go deep into code and architecture when needed, while giving engineers and technical leads the clarity and ownership to do their best work. They set technical direction, grow strong teams, make durable architecture decisions, and turn ambiguous 0-to-1 problems into platforms OpenAI can trust and build on for years. In this role, you will: Build and lead a high-performing Identity Infrastructure team, going deep enough technically to set direction while empowering the team to own delivery. Define the strategy for identity platform as the policy plane for access across people, agents, workloads, services, clouds, and internal systems. Scale Acc

awsgitrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

awskubernetesci/cd
View job →

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer you will join our OpenAI engineers and researchers in building, operating and securing transformational AI technologies. This role will focus on all aspects of Detection & Response but with a strong emphasis on detecting insider threats and influencing controls to safeguard OpenAI's most sensitive assets. In this role, you will: In this role, you will: Innovate on Detection and Response infrastructure to engineer and automate end-to-end detection and investigation workflows. Develop, measure, and tune detection rules to ensure effective and sustainable operations. Drive projects across OpenAI’s technology stack with a focus on insider threats, ranging from access abuse and intellectual property theft to novel risks emerging within AI infrastructure. Partner closely with cross-functional stakeholders, including HR, Legal, and peer investigative teams, providing technical expertise and evidence to support investigations. Collaborate on cutting-edge AI research, and use AI to improve OpenAI’s Security posture. You might thrive in this role if you: 5+ years experience working in a detection/response or insider-risk role.. We are seeking mid-level and senior candidates. You have broad familiarity with operating systems and platforms such as macOS, Windows, Linux, and Kubernetes, along with experience in cloud infrastructure. Knowledge of modern adversary tactics and attack paths, data exfiltration techniques, and h

pythonawskubernetes
View job →

About us Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. Role Overview We are seeking a Database STL (Individual Contributor) with deep expertise in MySQL and strong working knowledge of MongoDB, PostgreSQL, and Cassandra. This role combines hands-on database administration and optimization with strategic ownership of database reliability, automation, and cloud adoption. The candidate will lead by example—driving technical excellence, influencing best practices, and partnering cross-functionally with DevOps, SRE, and product engineering teams to deliver highly available, secure, and scalable database platforms. Key Responsibilities 1. End-to-End Ownership of MySQL databases in production & staging—availability, performance, and reliability. 2. Architect, manage, and support MongoDB, PostgreSQL, and Cassandra clusters for scale and resilience. 3. Define and enforce backup, recovery, HA, and DR strategies across all critical database platforms. 4. Drive database performance engineering—tuning queries, optimizing schemas, indexing, and partitioning for high-volume workloads. 5. Own replication, clustering, and failover architectures ensuring business continuity. 6. Champion automation & AI-driven operations—design self-healing scripts, predictive scaling, and proactive monitoring solutions. Collaborate with Cloud/DevOps teams on AWS database services (RDS, Aurora, DynamoDB, EC2, S3) to optimize cost, security, and performance. 7. Establish monitoring dashboards & alerting mechanisms for slow queries, replication lag, deadlocks, and capacity planning. Ensure compliance & security standards—encryption, auditing, and regulatory requirements. 8. Lea

sqlpostgresqlmysql
View job →
TI
TextNow, Inc.
📍 San Francisco• Full-time• $136.3K – $273.9K/yr
15 days ago

TextNow is on a mission to make communications affordable and accessible for everyone. As a full MVNO operating our own mobile core network over LTE and 5G NSA, we have the unique advantage of controlling our network infrastructure end-to-end. We operate the HSS, PGW, and other critical network functions, giving us the flexibility to innovate and deliver exceptional service to millions of users. About the Role We are looking for a hands-on Tech Lead, Mobile Core Network Engineering to own the health and evolution of our mobile core platform. Our cloud-hosted core runs in AWS, with our mobile core vendor providing managed services for day-to-day operations. This role is the critical bridge between our vendor’s Managed Services team and our internal engineering, product, and operations teams. You’ll be the technical owner of the mobile core, combining deep technical expertise with strong coordination skills. You won’t just manage vendor deliverables—you’ll dive into the technical details, challenge approaches when needed, and ensure our core network evolves to meet the demands of a growing subscriber base. This role has a clear growth path into technical and people leadership as the team scales. A note on background: We know that people who have operated a mobile core at scale are rare. If you’re a strong SRE or infrastructure engineering leader with deep networking knowledge and a track record of managing complex, vendor-operated platforms, we want to hear from you—even if you haven’t worked with 3GPP technologies before. We’ll invest in helping you build the telco-specific domain knowledge. What You'll Do Own the Mobile Core Platform: Take end-to-end accountability for the health, performance, and evolution of TextNow's mobile core (HSS, PGW, PCRF, and related functions running in AWS). Drive capacity plann

pythonawskubernetes
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

We are seeking an Engineering Manager to join our growing Gurugram Product & Technology team to provide technical direction, direct architecture, and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As an Engineering Manager on this new team, you will be responsible for leading and growing an engineering team, taking on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement Drive the architectural vision for the platform, ensuring it runs equally well on public clouds, private cloud environments and on-premise Work with product managers, program managers, design & analytics teams and other teams to define, prioritize and deliver new features that delight our users and drive platform improvements Take responsibility for the planning and execution of major features, raise delivery risks Own the monitoring, operations, and maintenance of the systems your team develops Enable the team to operate efficiently by removing technical obstacles, coordinating with other teams on dependencies, and prioritizing the team's overall well-being Contribute to planning for organizational growth, including allocation of engineering resources, participate in hiring and assignment of projects Qualifications 8+ years of experience of building distributed systems, and/or foundatio

pythonjavamongodb
View job →

NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product workloads. The company seeks a Senior Technical Program Manager (TPM) to lead complex, cross-functional programs powering NVIDIA’s next-generation AI software platforms. In this role, you will drive software initiatives across platform services, cloud infrastructure, and system integration. The focus is on enabling scalable, reliable, and supportable software for AI workloads. You will be responsible for managing high-impact engineering programs within a dynamic, fast-paced roadmap, aligning priorities across teams, and ensuring timely, high-quality delivery. This role requires strong technical competence, a proactive approach, and the ability to operate effectively across multiple levels of the organization. This is a software-first TPM role. The ideal candidate has extensive experience managing software initiatives. They also understand the full-stack environment, including infrastructure dependencies, system bring-up, integration readiness, and operational needs to support software across stack layers. What You'll Be Doing: Lead end-to-end execution of software platform initiatives, including planning, execution, delivery, and operationalization. Work together with software, infrastructure, product, and operations teams to ensure alignment on goals, deliverables, achievements, and schedules. Lead cross-functional initiatives encompassing cloud-native services, platform software, system integration, and release delivery. Help connect software roadmap execution to full-stack readiness, including dependencies across infrastructure, bring-up, validation, and downstream operational support. Identify cross-functional dependencies, mitigate risks, and drive resolution of complex technical and programmatic issues. Establish clear success metrics and reporting mechanis

kuberneteslinuxmachine learning
View job →

AANZARA CORPORATE IS HIRING! Join Our Growing Team AANZARA Corporate is looking for passionate and career-driven professionals to join our expanding team across multiple departments. Freshers and experienced candidates are welcome to apply. Locations: Vallioor Tirunelveli Nagercoil Coimbatore FMCG Department - Channel Sales Executive (Field Sales) Salary: 30,000 40,000/month Job Responsibilities: Visit retailers, Consumers, and supermarkets. Generate new business and maintain customer relationships. Promote FMCG products and achieve monthly sales targets. Collect orders and ensure product availability. Submit daily sales and market reports. Eligibility: 12th / Any Degree Freshers & Experienced Bike & Valid Driving Licence Mandatory Good Communication Skills HR Department HR Recruiter (Office Based) HR Marketing Executive Location: Vallioor Salary: 20,000/month Job Responsibilities: Source and screen candidates. Schedule interviews and coordinate recruitment. Handle employee onboarding and HR documentation. Maintain recruitment reports and candidate database. Support HR operations and employee records. Eligibility: Any Degree / MBA Preferred Good Communication Skills Basic Computer Knowledge Freshers & Experienced Software Department Location: Vallioor Tirunelveli Salary: 20,000+ (Based on Skills & Experience) Open Positions Frontend Developer Cloud Engineer Figma UI/UX Designer Testing Engineer Skills Required: Frontend Developer HTML, CSS, JavaScript React.js / Next.js Responsive Web Design Git & API Integration Cloud Engineer AWS / Azure / Google Cloud Linux & Networking CI/CD Pipelines Docker & Kubernetes (Preferred) Figma UI/UX Designer Figma UI/UX Design Wireframes & Prototypes Design Systems Testing Engineer Manual & Automation Testing Test Case Preparation Bug Reporting STLC / SDLC Knowledge Benefits Attractive Salary Package Performance Incentives Career Growth Opportunities Professional Training Supportive Work Environment Long-Term Career

javascriptreactaws
View job →

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. Our team develops the operating model that connects infrastructure strategy, supply planning, manufacturing operations, and delivery into a single, integrated system that enables OpenAI to deploy AI infrastructure predictably at scale. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Strategic Sourcing, and external infrastructure partners to create a single, integrated view of program health. Through governance, operational analytics, executive reporting, and scalable operating mechanisms, we enable leaders to proactively manage risk, optimize capacity, and deliver infrastructure predictably at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive integrated strategy and delivery across OpenAI's rapidly expanding AI infrastructure portfolio. This role sits at the intersection of infrastructure strategy, New Product Introduction (NPI), supply planning, manufacturing operations, and infrastructure delivery. You will lead highly cross-functional programs spanning engineering, supply planning, manufacturing, logistics, construction, commissioning, and operations, ensuring technical and operational dependencies remain synchronized from planning through production readiness. Beyond driving program execution, you will leverage operational insights to improve capacity planning, infrastructure strategy, and deployment readiness. You will also help operationalize new technologies and suppliers by partnering w

REMOTEawsrestagile
View job →
M
Mongodb
📍 Bengaluru• Full-time
1mo ago

We are seeking an Engineering Manager to join our growing Gurugram Product & Technology team to provide technical direction, direct architecture, and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As an Engineering Manager on this new team, you will be responsible for leading and growing an engineering team, taking on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Position Expectations Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement Drive the architectural vision for the platform, ensuring it runs equally well on public clouds, private cloud environments and on-premise Work with product managers, program managers, design & analytics teams and other teams to define, prioritize and deliver new features that delight our users and drive platform improvements Take responsibility for the planning and execution of major features, raise delivery risks Own the monitoring, operations, and maintenance of the systems your team develops Enable the team to operate efficiently by removing technical obstacles, coordinating with other teams on dependencies, and prioritizing the team's overall well-being Contribute to planning for organizational growth, including allocation of engineering resources, participate in hiring and assignment of projects Qualifications 8+ years of experience of building distributed systems, and/or foundati

pythonjavamongodb
View job →
M
Mongodb
📍 Gurugram• Full-time
1mo ago

We are seeking an Engineering Lead to join our growing Gurugram Product & Technology team to provide technical direction, direct architecture, and implement core parts of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As an Engineering Lead on this new team, you will be responsible for leading and growing an engineering team, taking on challenging, high-visibility projects that improve and enhance the performance, scalability, and reliability of the distributed systems infrastructure for this new product. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Provide technical leadership and mentorship to a team of engineers, fostering a culture of innovation, quality, and continuous improvement Drive the architectural vision for the platform, ensuring it runs equally well on public clouds, private cloud environments and on-premise Work with product managers, program managers, design & analytics teams and other teams to define, prioritize and deliver new features that delight our users and drive platform improvements Take responsibility for the planning and execution of major features, raise delivery risks Own the monitoring, operations, and maintenance of the systems your team develops Enable the team to operate efficiently by removing technical obstacles, coordinating with other teams on dependencies, and prioritizing the team's overall well-being Contribute to planning for organizational growth, including allocation of engineering resources, participate in hiring and assignment of projects Qualifications 8+ years of experience of building distributed systems, and/or foundational cl

pythonjavamongodb
View job →
O
Okta
📍 Bellevue, Washington; Chicago, Illinois; New York, New York; San Francisco, California; Washington, DC• Full-time• From $194K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team The Site Reliability team is dedicated to architecting and owning the foundational infrastructure tooling and CI/CD platforms that support Okta’s SRE ecosystem. In this development-focused role, you will leverage a modern tech-stack to build durable, automated systems that maximize platform reliability and engineering velocity. The ideal candidate is someone who enjoys analyzing systems and identifying areas of opportunity to improve system performance, availability and capacity. They are part systems administrator, part network administrator, and part developer. What you’ll be doing Maintain a highly available cloud infrastructure edge for the Okta identity platform Automate AWS infrastructure with Terraform and/or Chef Evolve the system by introducing changes to improve efficiency, scalability, and velocity What you’ll bring to the role 8+ years of operations experience configuring, deploying, monitoring and troubleshooting applications and

pythonawsdocker
View job →

About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Order to Cash (OTC) team oversees the complete flow of commercial transactions from order intake and provisioning through billing, collections, and cash application — ensuring accuracy, compliance, and operational excellence in support of OpenAI’s mission to ensure artificial general intelligence benefits all of humanity. About the Role We are looking for a senior, hands-on operator to own Order Management and Billing execution across OpenAI’s cloud marketplace and partner ecosystem, including platforms such as AWS, GCP, Oracle Cloud, GovCloud, and future channels. This senior individual contributor role will translate partner requirements into scalable workflows and ensure launch readiness, accurate billing, and reliable daily execution. As a senior individual contributor within the Cloud Marketplaces team, you will own the end-to-end order-to-invoice lifecycle for your assigned portfolio. You will ensure that private offers, commercial terms, provisioning, pricing, usage, billing data, credits, settlements, and partner-specific reporting flow through our systems accurately, on time, and with audit-ready controls. You will implement and continuously improve the common cloud marketplace operating model, lead cross-functional execution for your assigned portfolio, and surface risks, requirements, and improvement opportunities. You will partner across Revenue Systems, Product, Engineering, GTM, Finance, Partner Operations, and external marketplace stakeholders. This role is critical to building the operational backbone for OpenAI’s expansion across cloud marketplaces and government-cloud channels. You will combine deep operational judgment with process and control execution, automation, clear communication, and hands-on problem solving to improve billing reliability, partner experience, customer outcomes, and financial integrity at scale. This role

awsgcprest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s GTM Partnerships team builds a strategic, global partner ecosystem that accelerates customer success, enables enterprise AI adoption, and drives durable growth in support of OpenAI’s mission. Cloud partners are central to how customers discover, purchase, deploy, and scale OpenAI solutions. We work across cloud providers and their field, product, marketplace, technical, and partner organizations to reduce friction for customers and create repeatable paths from initial interest to production impact. The team collaborates closely with Sales, Technical Success, Product, Engineering, Finance, Legal, Security, Marketing, Operations, and Customer Success to turn strategic cloud relationships into measurable customer and commercial outcomes. About the Role We are hiring a Partner Director, GTM Cloud Partnerships to shape and execute OpenAI’s go-to-market strategy with strategic cloud partners. You will own senior relationships across a portfolio of cloud providers and develop joint business plans that expand customer access to OpenAI’s products, generate partner-sourced opportunities, and accelerate enterprise adoption. You will translate company-level partnerships into practical field motions spanning account planning, co-selling, cloud marketplaces, customer procurement, technical enablement, solution development, and executive engagement. This is a highly strategic and hands-on role. Success requires executive presence, strong commercial judgment, cloud ecosystem expertise, and the operating discipline to coordinate complex initiatives across multiple partners, regions, and internal teams. The ideal candidate can move fluidly between long-term strategy and the detailed execution required to deliver measurable results. This role is based in San Francisco. We use a hybrid work model of three days in the office per week. In this role, you will: Own the strategic and commercial relationships with a top global cloud partner, establishing strong alignm

awsrestai
View job →

About the Team OpenAI’s Governance, Risk, and Compliance team helps ensure security and privacy are grounded in how our products and systems actually operate. Assurance Operations partners with Security, Engineering, Infrastructure, Product, Privacy, and Legal to make controls provable, risk decisions explicit, and audit readiness a result of well-designed systems. About the Role We are hiring a technical, product-minded GRC builder who can own consequential audits while improving the control and evidence systems behind them. You will build a reusable common control framework, use Codex to automate assurance work, validate changing system scope, and turn repeated audit friction into measurable improvements. We are looking for someone who questions inherited assumptions, solves novel problems creatively, works closely with engineers, and makes the next audit easier by improving the underlying system. You’ll be responsible for: Lead external, internal, customer, and certification audit work from scoping through evidence review, fieldwork, remediation, and closeout. Build a common control framework linking risk, control intent, implementation, owner, system, environment, evidence, and applicable frameworks. Validate actual scope and ownership instead of assuming last year's controls, product boundaries, or evidence remain accurate. Use Codex to build and test evidence checks, control mappings, request triage, owner workflows, monitoring, and remediation reporting. Partner with engineers on cloud architecture, identity, logging, data flows, software changes, vulnerabilities, and control effectiveness. Design maintainable, permission-aware tools that preserve source provenance, human review, and evidence integrity. Reduce repeated requests and operational burden for control owners through measurable workflow improvements. Define roadmaps, decision rights, milestones, success metrics, and clear cross-functional escalations. We’re looking for someone with: Direct ownership

sqlawsrest
View job →
🔔

Get new cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More cloud operations engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.