Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: As a New Grad Software Engineer, you'll join a team of exceptional builders working on products that are reshaping how the world creates software. You'll have the opportunity to work on everything from our AI-powered development platform to the distributed systems that enable real-time collaboration for millions of developers. This is a chance to define your career while defining the future of software development. You'll work on problems that matter, with the autonomy to drive solutions and the support to grow into a technical leader. What you will build: Product features that delight users and make it possible for anybody to create software AI coding agent that understands intent and generates production-ready applications Cloud infrastructure that provides instant, powerful development environments at global scale Platform features that enable one click deployments and scale to millions of users Required skills and experience: Recent graduate (2027) with a degree in Computer Science, Computer Engineering, or related field Strong programming skills in a modern language (JavaScript/TypeScript, Python, Go, Rust) Full-stack capabilities with experience in React, Node.js, and database technologies Growth orientation - eager to learn new technologies and take on increasing responsibility Collaborative spirit - you work well in cross-functional teams and value diverse perspectives What we value : Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions Self-directed and autonomous: Capable of working independently while collaborating effectively with cross-functional teams Strong communication skills: Ability to explain complex technical conce
Jobs in United States
Cloud Operations Engineer in United States
698 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Staff Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A STAFF SOFTWARE ENGINEER, YOU WILL: Design and develop features, understand customer requirements and meet business goals. Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. Build highly reliable and fault-tolerant software to meet the needs of the largest customers. Ensure operational readiness and maintainabilit
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally. What You Will Be Doing: Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software. Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can b
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence. As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization. If you're p
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Infrastructure Software Engineer at Baseten, you'll build and maintain components of our ML inference platform that powers production AI applications. You'll contribute to the core infrastructure, enabling developers to deploy, scale, and monitor ML models with high performance. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Develop infrastructure components for our ML inference platform using Python and Go Implement and maintain Kubernetes deployments for model serving Contribute to our inference orchestration layer for model deployments Build and enhance monitoring systems for model performance metrics Implement efficient resource management solutions for ML workloads Support infrastructure automation to improve ML deployment workflows Work closely with team members to implement technical solutions Help balance performance optimization with system reliability Participate in technical discussions around infrastructure improvements Learn and apply infrastructure best practices REQUIREMENTS Bachelor's degree or higher in Computer Science or related field Proficient coding abilities in one or more popular programming or scripting languages; Go proficiency is a plus Working knowledge of Kubernetes and containeriza
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At Modal, we sell cloud services atop which our customers run their critical production systems. As a rapidly growing new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform, customer base, and our team. This role is for people who are deep systems thinkers, love stacking nines, and thrive from making others move faster at scale. Responsibilities include: Identifying architectural changes to improve reliability and performance. Fostering a culture of reliability across Modal’s engineering organization. Defining and implementing operational processes such as deployments, upgrades, etc. Operating systems like Kubernetes, Postgres, Redis, etc. Participating in on-call rotations, and responding to production incidents. Requirements: 5+ years of experience writing high-quality production code. 2+ years of
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a customer-obsessed software engineer to come ship with us. You’ll own features like multi-node training and products like serverless reinforcement learning (RL) from conception to MVP (and from MVP to GA!). You’ll work through the stack, architecting solutions from API and UI down to our infrastructure layer. You’ll fine tune models yourself to develop an understanding of user workflows. You’ll work closely with research engineers leveraging state-of-the-art training techniques to build experiences that accelerate model development and solve for real pain points. If you’re excited to dive deep into the training, let’s talk! THE PRODUCT Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done EXAMPLE INITIATIVES Checkpointing Pipeline: Our checkpointing pipeline starts with automated checkpointing, a feature that ensures that versions of models created during training are automatically backed up to the cloud. Users are able to then deploy checkpoints seamlessly into inference servers, providing point-and-click integrations into inference frameworks like vLLM and Baseten’s Inference Stack. This enables customers to quickly evaluate the performance of their checkpoints with real traffic. Multinode training: Multinode training enables customers to easily run training jobs across multiple compute nodes, enablin
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 We are seeking a highly technical and experienced Engineering Manager to one of our business critical engineering team. This role is ideal for a hands-on leader who thrives in a fast-paced environment, has a deep understanding of collaborative document editing technologies, and is passionate about building scalable, high-performance software. As the Manager, you will oversee the development and delivery of innovative features, ensure technical excellence, and mentor a team of talented engineers. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our product. Operational Excellence : Ensure the stability and performance of the Docs platform, proactively addressing technical debt and optimizing system architecture. Qualifications: Technical Expertise : Proficiency in
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our engineering team for a 12-week paid internship where you'll work alongside world-class engineers, designers, and product managers to build the future of software creation. You'll contribute to real features that impact millions of developers worldwide, from our AI-powered development environment to the infrastructure that makes lightning-fast collaboration possible. This isn't just about learning—you'll ship meaningful code that helps democratize software creation. Whether you're optimizing our cloud infrastructure, building intuitive developer tools, or enhancing our AI agents, your work will directly empower creators around the globe. You will: Ship real features to millions of developers using Replit's platform Collaborate cross-functionally with engineers, designers, product managers, and AI researchers Build and optimize developer experiences that make coding accessible to everyone Work on cutting-edge AI tools and infrastructure that power the next generation of software creation Learn from the best in an environment where your ideas are heard and often implemented Required skills and experience: Currently pursuing a Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or related technical field Have at least one semester of schooling remaining after the internship completion Proficient in at least one programming language and comfortable with full stack development Passionate about developer tools, AI, or making technology more accessible Thrive in fast-paced environments where you can move quickly and adapt to changing priorities What we value : Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions Se
$170K – $250K/yr
At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You'll Play At Mindbody, Core Engineering builds and evolves the foundational systems that help our products run reliably at scale. In this staff role, you’ll bring clarity to complex technical problems, guide architecture, and strengthen how we design, deliver, and operate the backend services that power real-world experiences. Lead cross-team technical execution, aligning architecture and delivery across multiple squads and core domains Design and evolve microservices patterns that improve reliability, performance, and maintainability Drive cloud and deployment improvements across AWS, Mindbody’s cloud platform, and our containerized deployment environment Partner with engineering and product leaders to turn ambiguous problems into clear technical plans and milestones Establish and socialize standards for service design, APIs, and relational data modeling (SQL) Strengthen monitoring and operational visibility using New Relic and Kibana, turning insights into durable system improvements Mentor and unblock engineers through design reviews, pairing, and practical guidance Reduce technical risk and complexity while balancing
Other cities to consider
More places hiring for this role
Get new cloud operations engineer jobs in United States by email
Daily job updates · Unsubscribe anytime