About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a VP of Finance to build the finance function from the ground up as our first full-time finance hire. This is a high-impact role for someone who thrives at the intersection of strategic thinking and hands-on execution. We are looking for someone who can architect the systems and processes that will scale with Modal, partner closely with the founders and executive team, and grow into the company's CFO. You'll report directly to the CEO and collaborate closely with our BizOps, GTM, and Product teams. In this role, you will: Build and maintain Modal's operating model, tying financial performance to company KPIs and resource allocation Lead all budgeting, forecasting, and long-range planning processes, and develop the reporting infrastructure that gives leadership and the board clear, timely visibility into the health of the business Partner with the found
Jobiba hiring network
Cloud Operations System Administrator Jobs
2,288 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity New Relic is looking for a Senior Business Value Engineer to join our Value Strategy team, specifically covering the EMEA region . In this role, you will be the primary value architect for our accounts across Europe, the Middle East, and Africa. You will bridge the gap between technical capabilities and business outcomes, helping our customers understand the financial and operational impact of the New Relic platform. You will act as a trusted advisor to both our internal sales teams and our customers' C-suite executives, driving the "Value Realization" process from initial discovery to long-term success. What you'll do Build Compelling Business Cases: Create and deliver Business Value Assessments (BVAs) including complex ROI modeling, TCO analysis, and value realization benchmarking. Strategize with Sales: Partner closely with Sales Leadership and Account Executives to develop deal strategies that emphasize business justification over feature-set comparisons. Quantify Technical Impact: Translate technical metrics (like MTTR, Error Rates, and Cloud Cost) into business KPIs (like Revenue at Risk, Subscriber Churn, and Operational Efficiency). Drive Value Realization: Ensure customers are co-creators in their value roadmap, moving beyond the initial sale to track and report on actual value achieved post-implementation. Educate & Enable: Mentor and train the broader GTM teams on value-based selling best practices to inc
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Platform Engineer, you will contribute to building and improving Stitch Fix’s cloud-native infrastructure and internal developer tooling. You’ll work on tools and automation that help product engineers deploy, operate, and debug services more easily, while learning modern platform engineering practices alongside experienced teammates. This role is ideal for engineers who enjoy improving developer experience and want to grow their skills in cloud infrastructure and CI/CD systems. Responsibilities: Contribute to the development and evolution of our internal platform-as-a-service used by application and service developers Build and maintain tooling that improves developer workflows, deployment reliability, and day-to-day productivity Collaborate with platform and application engineers to identify friction points and implement incremental improvements Learn and apply best practices around Infrastructure-as-Code, containerized workloads, and CI/CD pipelines Use, or are eager to adopt, AI-assisted development tools to improve productivity, and are excited to help explore and integrate LLM-powered solutions that automate internal support and operational workflows Have opportunities to propose ideas and improvements, with support and mentorship from the team Things you’ll get exposure to (and we don’t expect experience with everything): AWS Terraform, Pulumi CircleCI Docker, ECS, EKS Ruby, Golang, Python About You 2+ years of software development and infrastructure experience with significant contribut
As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i
As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i
We’re looking for an Engineering Manager to lead our Sensitive Data Scanner (SDS) Telemetry team. The SDS group’s mission is to be the world’s easiest-to-use tool to discover, classify, manage, and report sensitive data risks across cloud, on-premise, and code environments. This team builds and scales the detection capabilities that scan all telemetry data flowing into Datadog — logs, APM spans, and RUM events — operating in streaming, at processing time, and at very large scale. You’ll lead a small, close-knit team based in Paris, with the opportunity to shape how the team grows as SDS Telemetry’s scope expands. It’s a chance to combine hands-on technical leadership with direct customer and product impact in the security and observability space. At Datadog, we place value in our office culture — the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead and grow a team of engineers building real-time sensitive data detection across Datadog’s Logs, APM, and RUM telemetry pipelines Partner closely with the Logs, APM, and RUM teams, plus Datadog’s Trust & Safety team, to align on roadmap and integration priorities Shape product direction by working closely with Product, grounding decisions in customer needs and business impact Stay hands-on: contribute to design decisions and participate in the team’s on-call rotation Recruit, mentor, and develop engineers as the team grows beyond its initial size Help build a strong engineering culture as part of Datadog’s broader Sensitive Data Scanner group Who You Are: You have experience building and shipping revenue-generating products, with strong product acumen and a customer-first mindset You have hands-on experience with Go and/or Java, and a track record building distributed, streaming systems at scale You have experience managing engineers — or are
MongoDB Technical Services Engineers for Named Accounts, as part of the Premium Services team within Technical Services, will use their exceptional problem solving and customer service skills, along with their deep technical experience, to advise customers and solve their complex MongoDB problems. They are experts in one or more components of the MongoDB ecosystem - database server, drivers, our management suite, and services such as Cloud Manager (the online product we developed for customers for automation, backup, monitoring, and analysis of their MongoDB systems), and MongoDB Atlas. Our engineers combine their MongoDB expertise with passion, initiative, teamwork, and a great sense of humor to achieve exceptional results for our customers. We're looking to speak with candidates based in San Francisco or Palo Alto for our hybrid working model. Cool things you’ll do MongoDB is on a mission to change the way people think about databases. Along the way, our customers encounter questions and issues about how our approach to databases works for their use cases. In Technical Services, it's our job to help these people. You'll be primarily working alongside a handful of our largest enterprise customers - building a relationship, intimately guiding their use of MongoDB products and services, coordinating the resolution of their complex issues, - answering questions on architecture, performance, recovery, security, and everything in between. You'll be an expert resource on best practices for running MongoDB at scale, whatever that scale may be. You'll be an advocate for customers' needs, - interfacing with our product management and development teams on their behalf. What you need You should have 7+ years of database industry experience deploying and managing operational production databases at scale, both on premise and in the cloud. We encourage you to apply even if you’ve never used MongoDB before. We consider all candidates with an eye for those who are sel
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! Figma's AI Tools team builds the AI-powered workflows, platforms, and tooling that make every engineer at the company more productive. From agentic CI auto-fixing and AI-assisted code review to background cloud agents that turn a Slack message into a pull request, this team owns the systems that are transforming how Figma builds software. The team operates across a three-layer platform stack—sandbox runtime, cloud agents, and workflow orchestration—while shipping and maintaining a growing portfolio of org-wide developer workflows that compound across hundreds of engineers. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you’ll do at Figma: Lead and grow a team of engineers responsible for building and operating Figma's AI developer workflows and the cloud agent platform that powers them Own the technical strategy and roadmap for AI Developer Experience, spanning sandbox runtime infrastructure, cloud agent reliability, workflow orchestration, and org-wide agentic workflows Hire and scale the team - establishing team culture, execution cadence, and operational processes from the ground up Drive the reliability, observability, and scalability of our cloud agent platform, ensuring it meets production-grade standards as adoption grows across the engineering organization Partner with product engineering, security, infrastructure, and DevEx teams to identify the highest-leverage opportunities for AI-assisted developer workflows and drive adoption
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We’re seeking an exceptional Principal-level Offensive Security Engineer focused on deep, hands-on penetration testing of OpenAI’s agent-powered products, infrastructure, and model-integrated application surfaces. You’ll assess complex systems end to end, identify realistic vulnerabilities, validate exploitability and impact, and partner closely with engineering teams to drive durable fixes. This role will be primarily focused on continuously testing our agent-powered products like Codex and Operator. These systems are uniquely valuable targets because they’re rapidly evolving, can perform sensitive actions on behalf of users, and have large, diverse attack surfaces. You will play a crucial role in securing our agents by finding vulnerabilities that emerge from the interactions between the applications, infrastructure, tools, and models that power them. You’ll have the chance to not only find vulnerabilities, but actively drive their resolution, build reusable testing approaches, automate offensive security workflows with cutting-edge technologies, and use your attacker perspective to improve the security of OpenAI’s products. In this role you will: Conduct deep penetration tests of OpenAI’s agent-powered products, including web applications, APIs, cloud services, identity and authorization flows, CI/CD systems, and model-integrated product surfaces. Continuously hunt for exploitable vulnerabilities in the interactions between the appli
About the Role We are looking for a Director of Engineering – AI & Full-Stack SaaS to lead the engineering strategy and execution for a highly scalable, enterprise-grade SaaS platform. This role is ideal for a strong engineering leader who combines deep software engineering expertise, enterprise SaaS experience, cloud-native architecture, full-stack product development, and practical AI/GenAI adoption . The successful candidate will lead multiple engineering teams, drive architectural and technical decisions, improve engineering velocity and quality, and help embed AI across the software development lifecycle and product engineering ecosystem. Key Responsibilities Lead and mentor multiple engineering teams responsible for building and delivering enterprise SaaS products. Define and execute the engineering strategy, technical roadmap, architecture and development standards . Drive development of highly scalable, secure, reliable and cloud-native SaaS applications. Provide technical leadership across backend, frontend, APIs, microservices and distributed systems . Drive adoption of AI/GenAI tools and capabilities across the engineering lifecycle, including development, testing, code quality, productivity and automation. Partner closely with Product, Architecture, DevOps, Security and other cross-functional teams to deliver business-critical capabilities. Establish engineering best practices around design, coding, testing, CI/CD, observability, performance and reliability . Own engineering delivery, quality, scalability and operational excellence across multiple product areas. Identify and resolve complex technical and architectural challenges. Build a high-performing engineering culture focused on innovation, ownership, collaboration and continuous improvement . Evaluate emerging technologies and identify opportunities to leverage AI and automation to improve engineering productivity and product capabilities. Drive technical modernization and evolution of existing
Here's a summary of the role: You're a hands-on engineer who enjoys solving complex data challenges at scale. You'll join the Data Hub team, building the shared data platform that supports reporting, analytics, AI, and Data Science initiatives across Diligent. Working with Python, SQL, AWS, and cloud-native technologies, you'll help process, organize, and enrich large volumes of data while developing scalable services and modern data solutions. You'll collaborate closely with Product Managers, Engineering Managers, and fellow engineers to deliver high-quality features, improve platform capabilities, and shape the future of our data ecosystem. If you enjoy ownership, solving challenging problems, experimenting with new technologies, and working in a collaborative environment, you'll fit right in. Here's a breakdown of what you'll do (not all of it, just the important stuff): Design, develop, review, and test features and user stories following Agile development practices Contribute to core platform development and integration projects across multiple systems Participate in shaping the future architecture of the product , including designing scalable backend services and prototypes Collaborate with Product Owners and Engineering Managers to analyze, refine, and document technical requirements Build and maintain cloud-native solutions and microservices-based applications Leverage AI-assisted development tools , code assistants, and modern engineering workflows to improve delivery efficiency Support the creation and maintenance of high-quality operational and technical documentation Continuously identify opportunities to improve engineering processes, quality, and team effectiveness These are the essentials you'll need to get an interview: 2+ years of experience in a hands-on software development role within a commercial software environment Strong programming experi
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h
We're looking for an ML Data & Platform Engineer to own the infrastructure that powers our speech AI models: the pipelines that source and prepare training data, and the platform that trains, evaluates, and serves them in production. Speech AI has a data problem most ML teams don't, and you'll be at the centre of solving it, working as part of our ML team to remove friction across the entire lifecycle and get better models into production faster. This is a broad, cross-functional role suited to someone who enjoys working across the full stack: data infrastructure, distributed systems, and production ML, and who takes ownership of problems end to end rather than waiting to be told what to fix. What you'll do Designing, building, and maintaining scalable data pipelines for ingesting, transforming, validating, and storing large datasets used to train our models Developing and maintaining web scraping and data acquisition solutions to keep training datasets fresh, high-quality, and available at scale Building and operating the infrastructure that lets the ML team deploy and evaluate new models quickly, and that serves models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning data infrastructure, training, and inference Continuously improving our data and MLOps practices, and helping shape the roadmap for how our platform evolves as we scale What you'll need Strong proficiency in Python and SQL, with a solid backend or data engineering foundation Hands-on experience with containerisation and orchestration (Docker, Kubernetes), and working with a major cloud provider Experience building data pipelines and ETL/ELT processe
Get new cloud operations system administrator jobs by email
Daily job updates · Unsubscribe anytime