Jobiba hiring network

Cloud Operations Engineer Jobs

2,288 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a customer-obsessed software engineer to come ship with us. You’ll own features like multi-node training and products like serverless reinforcement learning (RL) from conception to MVP (and from MVP to GA!). You’ll work through the stack, architecting solutions from API and UI down to our infrastructure layer. You’ll fine tune models yourself to develop an understanding of user workflows. You’ll work closely with research engineers leveraging state-of-the-art training techniques to build experiences that accelerate model development and solve for real pain points. If you’re excited to dive deep into the training, let’s talk! THE PRODUCT Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done EXAMPLE INITIATIVES Checkpointing Pipeline: Our checkpointing pipeline starts with automated checkpointing, a feature that ensures that versions of models created during training are automatically backed up to the cloud. Users are able to then deploy checkpoints seamlessly into inference servers, providing point-and-click integrations into inference frameworks like vLLM and Baseten’s Inference Stack. This enables customers to quickly evaluate the performance of their checkpoints with real traffic. Multinode training: Multinode training enables customers to easily run training jobs across multiple compute nodes, enablin

kubernetesrestmachine learning
View job →
C
1mo ago

The Role We are looking for a Senior Partner Deployed Engineer to join the Customer Solutions team, focused on building Coder’s partner ecosystem across EMEA. This is a new function at Coder, modeled on the Forward Deployed Engineer role pioneered by companies like Palantir and now the fastest-growing technical role at frontier AI companies. The difference: instead of embedding with a single customer to deploy a platform, you embed with strategic partners to help them understand, position, and deliver Coder’s platform across their entire customer base. You are part Field CTO, part Industry Strategist, part Technical Specialist, and part Partner Relations Lead. You will be the technical authority within our EMEA partner ecosystem: setting the vision for how Coder fits into each partner’s AI, cloud, and modernization offerings, co-creating the GTM sales plays that partner sellers take to market, building the demonstrations and workshops that generate pipeline, and producing the Partner Relations content that establishes Coder’s technical brand in the ecosystem. This is not a support role. You are equally comfortable holding a strategic roadmap conversation with a partner CTO and debugging a Kubernetes deployment in a partner’s lab in the same afternoon. You are energized by the challenge of building a new category through partnerships, motivated by the multiplied impact of enabling an entire ecosystem rather than a single customer, and capable of operating with full autonomy in a fast-moving startup environment. What You’ll Do Serve as the strategic technical thought partner to EMEA partner leadership, including practice leads, CTOs, and solutions architects at global systems integrators, regional cloud consultancies, and hyperscaler field teams. Own the technical relationship and set the vision for how Coder fits into each partner’s AI and modernization portfolio. Co-create sales plays tailored to each partner’s customer base, vertical focus, and services capabilitie

awsazurekubernetes
View job →
C
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc

awsazuregcp
View job →
C
Clickup
📍 Canada• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ

node.jsawsmachine learning
View job →
C
Clickup
📍 United States• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ

node.jsawsmachine learning
View job →
C
Clickup
📍 United States• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 We are seeking a highly technical and experienced Engineering Manager to one of our business critical engineering team. This role is ideal for a hands-on leader who thrives in a fast-paced environment, has a deep understanding of collaborative document editing technologies, and is passionate about building scalable, high-performance software. As the Manager, you will oversee the development and delivery of innovative features, ensure technical excellence, and mentor a team of talented engineers. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our product. Operational Excellence : Ensure the stability and performance of the Docs platform, proactively addressing technical debt and optimizing system architecture. Qualifications: Technical Expertise : Proficiency in

node.jsawsmachine learning
View job →
R
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our engineering team for a 12-week paid internship where you'll work alongside world-class engineers, designers, and product managers to build the future of software creation. You'll contribute to real features that impact millions of developers worldwide, from our AI-powered development environment to the infrastructure that makes lightning-fast collaboration possible. This isn't just about learning—you'll ship meaningful code that helps democratize software creation. Whether you're optimizing our cloud infrastructure, building intuitive developer tools, or enhancing our AI agents, your work will directly empower creators around the globe. You will: Ship real features to millions of developers using Replit's platform Collaborate cross-functionally with engineers, designers, product managers, and AI researchers Build and optimize developer experiences that make coding accessible to everyone Work on cutting-edge AI tools and infrastructure that power the next generation of software creation Learn from the best in an environment where your ideas are heard and often implemented Required skills and experience: Currently pursuing a Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or related technical field Have at least one semester of schooling remaining after the internship completion Proficient in at least one programming language and comfortable with full stack development Passionate about developer tools, AI, or making technology more accessible Thrive in fast-paced environments where you can move quickly and adapt to changing priorities What we value : Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions Se

J
Justworks
📍 New York• Full-time• $167.5K – $235K/yr
1mo ago

Who We Are At Justworks, you’ll enjoy a welcoming and casual environment, great benefits, wellness program offerings, company retreats, and the ability to interact with and learn from leaders in the startup community. We work hard and care about our most prized asset - our people. We’re helping businesses get off the ground by enabling them to focus on running their business. We solve HR issues. We’re data-driven and never stop iterating. If you’d like to work in a supportive, entrepreneurial environment, are interested in building something meaningful and having fun while doing it, we’d love to hear from you. We're united by shared goals and shared motivations at Justworks. These are best summed up in our company values, which are reflected in our product and in our team. Our Values If this sounds like you, you’ll fit right in. Who You Are Justworks is looking for an experienced security engineer skilled in detection and response, who can help enhance and mature Justworks’ Security. As a Senior Detection Engineer, you’ll design, build, and maintain the detection logic that powers our platform, conduct proactive threat hunting, and drive continuous improvements across our detection and incident handling workflows. You’ll collaborate closely with IT, Engineering, Platform, and other members of the Security team to identify attacker behaviors, build high‑fidelity detections, and strengthen our defenses. You’ll also play a key role in designing and conducting table‑top exercises, improving processes, and building automation that reduces friction and accelerates response. You’ll help explore how AI can enhance detection, hunting, and operational efficiency. Your Success Profile What You Will Work On Build, tune, and deploy high‑quality detections across our platform Develop and refine detections using telemetry from EDR, threat intel, endpoint & cloud posture platforms and native AWS cloud services Conduct proactive threat hunting to uncover threat actor behaviors a

awsrestai
View job →
M
Mindbody
📍 United States• Full-time• $170K – $250K/yr
1mo ago

At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You'll Play At Mindbody, Core Engineering builds and evolves the foundational systems that help our products run reliably at scale. In this staff role, you’ll bring clarity to complex technical problems, guide architecture, and strengthen how we design, deliver, and operate the backend services that power real-world experiences. Lead cross-team technical execution, aligning architecture and delivery across multiple squads and core domains Design and evolve microservices patterns that improve reliability, performance, and maintainability Drive cloud and deployment improvements across AWS, Mindbody’s cloud platform, and our containerized deployment environment Partner with engineering and product leaders to turn ambiguous problems into clear technical plans and milestones Establish and socialize standards for service design, APIs, and relational data modeling (SQL) Strengthen monitoring and operational visibility using New Relic and Kibana, turning insights into durable system improvements Mentor and unblock engineers through design reviews, pairing, and practical guidance Reduce technical risk and complexity while balancing

pythonreactsql
View job →
M
Mindbody
📍 Brazil• Full-time
1mo ago

At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You’ll Play As a Senior Platform Engineer on Playlist, you’ll design and deliver Infrastructure-as-Code solutions that empower developer teams. You’ll drive cloud architecture, iterate with squads on their workloads, and build self-service tools to speed delivery and improve quality. Your work will help design, implement, and operate the cloud infrastructure that powers the Mindbody ecosystem and supports millions of users. Partner with Product and Engineering to design, build, and operate the cloud platform that enables squads to deliver reliably and autonomously. Own and evolve our production Kubernetes platform and core cloud primitives, driving safe, automated, and observable infrastructure-as-code. Deliver self-service tooling and IaC patterns so teams can provision and run workloads without platform intervention. Lead cross-team projects from problem definition through architecture, implementation, and launch while reducing operational toil and improving reliability and security. Be the go-to engineer for production incident response, runbook automation, and platform change management across segmented and compliance-bound environments. Experience You Bring

typescriptpythonaws
View job →
M
Mindbody
📍 Brazil• Full-time
1mo ago

At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You’ll Play As a Senior Platform Engineer on Playlist, you’ll design and deliver Infrastructure-as-Code solutions that empower developer teams. You’ll drive cloud architecture, iterate with squads on their workloads, and build self-service tools to speed delivery and improve quality. Your work will help design, implement, and operate the cloud infrastructure that powers the Mindbody ecosystem and supports millions of users. Partner with Product and Engineering to design, build, and operate the cloud platform that enables squads to deliver reliably and autonomously. Own and evolve our production Kubernetes platform and core cloud primitives, driving safe, automated, and observable infrastructure-as-code. Deliver self-service tooling and IaC patterns so teams can provision and run workloads without platform intervention. Lead cross-team projects from problem definition through architecture, implementation, and launch while reducing operational toil and improving reliability and security. Be the go-to engineer for production incident response, runbook automation, and platform change management across segmented and compliance-bound environments. Experience You Bring Senior

typescriptpythonaws
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The ability to monitor and assist our vehicles remotely plays a key role in our business strategy. As a Software Engineer, Video Streaming you will work on our in-house Teleoperations platform. You will work with a diverse team of engineers to build the core communication system as well as the cloud platform to connect vehicles and operators. This position involves broad technical understanding in networking algorithms, bandwidth estimation, rate control, computer networking, and real-time communication systems. The team is expected to deliver reliable solutions and license to 3rd party teleoperation usages. About the Work Design and implement an efficient pipeline with state-of-the-art video streaming techniques to deliver high priority real-time data stream Build an offline streaming simulation/emulation framework that can help to iterate the video streaming algorithm and predict online performance Test systems in real-w

SF
Stitch Fix
📍 United States• Full-time• Remote• From $88.1K/yr
1mo ago

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Platform Engineer, you will contribute to building and improving Stitch Fix’s cloud-native infrastructure and internal developer tooling. You’ll work on tools and automation that help product engineers deploy, operate, and debug services more easily, while learning modern platform engineering practices alongside experienced teammates. This role is ideal for engineers who enjoy improving developer experience and want to grow their skills in cloud infrastructure and CI/CD systems. Responsibilities: Contribute to the development and evolution of our internal platform-as-a-service used by application and service developers Build and maintain tooling that improves developer workflows, deployment reliability, and day-to-day productivity Collaborate with platform and application engineers to identify friction points and implement incremental improvements Learn and apply best practices around Infrastructure-as-Code, containerized workloads, and CI/CD pipelines Use, or are eager to adopt, AI-assisted development tools to improve productivity, and are excited to help explore and integrate LLM-powered solutions that automate internal support and operational workflows Have opportunities to propose ideas and improvements, with support and mentorship from the team Things you’ll get exposure to (and we don’t expect experience with everything): AWS Terraform, Pulumi CircleCI Docker, ECS, EKS Ruby, Golang, Python About You 2+ years of software development and infrastructure experience with significant contribut

REMOTEpythonawsdocker
View job →
A
Asana
📍 New York• Full-time• $248K – $282K/yr
1mo ago

We are looking for a deeply technical, customer-facing Forward Deployed Engineer to help product development organizations deploy Command - a new offering from Asana. As we incubate new motions and capabilities on Command, we need engineers who can work directly with customers, move quickly from insight to implementation, and bring real-world learning back into the product. This role is based in our San Francisco or New York City office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Collaborate with Engineering leaders, platform teams, administrators, and hands-on builders to connect Command to the systems where work already happens Establish the technical and operational foundations for adoption and help teams develop more effective ways of planning, building, shipping, and learning Move between cloud infrastructure, integrations, workflow design, automation, product enablement, and engineering coaching Work directly in customer environments — configuring secure access to source-control, CI/CD, ticketing, and identity systems, or pairing with engineering teams to design release workflows and help leaders clarify the context their agents and teams need to make good decisions Write production-quality code when needed and turn customer-specific solutions into reusable product capabilities Leave customers with more than a working deployment — leave them with a stronger system for building software About you 6+ years of experience in software engineering, forward deployed engineering, customer engineering, solutions engineering, implementation engineering, or a similar technical customer-facing role Strong coding and systems-thi

ci/cdrestai
View job →
A
Asana
📍 San Francisco• Full-time• $248K – $282K/yr
1mo ago

We are looking for a deeply technical, customer-facing Forward Deployed Engineer to help product development organizations deploy Command - a new offering from Asana. As we incubate new motions and capabilities on Command, we need engineers who can work directly with customers, move quickly from insight to implementation, and bring real-world learning back into the product. This role is based in our San Francisco or New York City office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Collaborate with Engineering leaders, platform teams, administrators, and hands-on builders to connect Command to the systems where work already happens Establish the technical and operational foundations for adoption and help teams develop more effective ways of planning, building, shipping, and learning Move between cloud infrastructure, integrations, workflow design, automation, product enablement, and engineering coaching Work directly in customer environments — configuring secure access to source-control, CI/CD, ticketing, and identity systems, or pairing with engineering teams to design release workflows and help leaders clarify the context their agents and teams need to make good decisions Write production-quality code when needed and turn customer-specific solutions into reusable product capabilities Leave customers with more than a working deployment — leave them with a stronger system for building software About you 6+ years of experience in software engineering, forward deployed engineering, customer engineering, solutions engineering, implementation engineering, or a similar technical customer-facing role Strong coding and systems-thi

ci/cdrestai
View job →
🔔

Get new cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More cloud operations engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.