At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build Core Data Engineering Primitives at Cloud Scale Data pipelines are foundational infrastructure — when they're fast, correct, and maintainable, customers build on them with confidence. If you've spent the bulk of your career building large-scale data infrastructure — designing streaming or transformation primitives, reasoning hard about consistency and fault tolerance, and owning the systems that run under millions of customer workloads — this role might be for you. You'll be working on the streaming and transformation layer at Snowflake: the constructs that define how customers move, shape, and maintain data. AI has a real presence in this work — in how customers use these pipelines and in how we think about building them — but the core job is hard distributed systems engineering, and that's what we're hiring for. About the Team We build the core data engineering primitives that power Snowflake's streaming and transformation capabilities. From the constructs customers use to define real-time pipelines to the execution fabric that makes those pipelines reliable and cost-efficient at cloud scale, our team owns the full stack of declarative data engineering. We're a small, high-ownership team operating close to the product — which means your decisions ship, your architec
Jobiba hiring network
Cloud Operations System Administrator Jobs
2,329 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team OpenAI’s Stargate and 3P Engineering teams are responsible for building and scaling the external infrastructure ecosystem that powers advanced AI systems. We work across hyperscalers, colocation providers, cloud partners, and strategic third-party operators to turn contracted capacity into production-ready compute. Our scope spans the full lifecycle of external deployments: commercial alignment, technical readiness, network integration, hardware enablement, operational readiness, and long-range scaling strategy. As OpenAI’s infrastructure footprint expands globally, we need leaders who can convert complex partner environments into reliable, high-velocity capacity for training and inference workloads. About the Role We are seeking a Technical Program Manager, Token-as-a-Service (TaaS) to lead delivery of external compute capacity that directly serves OpenAI model workloads. In this role, you will own complex cross-functional programs that transform third-party infrastructure into usable tokens at scale. You will partner across engineering, capacity planning, networking, hardware, finance, product, and external providers to ensure that deployed capacity translates into real production throughput. This role sits at the intersection of infrastructure execution, systems readiness, and business impact. Success requires strong technical fluency, elite program management, and the ability to drive accountability across internal teams and external partners. This is a high-visibility role with direct impact on OpenAI’s ability to scale model training and inference globally. This role is based in San Francisco, CA, with a hybrid work model of 3 days in office per week. Relocation assistance is available. Key Responsibilities Lead end-to-end delivery programs that convert external infrastructure capacity into production-ready token supply. Own readiness across compute, storage, networking, security, and operational dependencies for third-party environments. Build
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build Core Data Engineering Primitives at Cloud Scale Data pipelines are foundational infrastructure — when they're fast, correct, and maintainable, customers build on them with confidence. If you've spent the bulk of your career building large-scale data infrastructure — designing streaming or transformation primitives, reasoning hard about consistency and fault tolerance, and owning the systems that run under millions of customer workloads — this role might be for you. You'll be working on the streaming and transformation layer at Snowflake: the constructs that define how customers move, shape, and maintain data. AI has a real presence in this work — in how customers use these pipelines and in how we think about building them — but the core job is hard distributed systems engineering, and that's what we're hiring for. About the Team We build the core data engineering primitives that power Snowflake's streaming and transformation capabilities. From the constructs customers use to define real-time pipelines to the execution fabric that makes those pipelines reliable and cost-efficient at cloud scale, our team owns the full stack of declarative data engineering. We're a small, high-ownership team operating close to the product — which means your decisions ship, your architec
NetSuite Developer The NetSuite Developer is responsible for the configuration and implementation of the brand-new NetSuite ERP solution based on business requirements and existing operational models. The ideal candidate will be responsible for configuration, programming and/or implementing as well as administering NetSuite application. Responsibilities Include: Assist with full NetSuite implementations as well as maintaining customer that have been using NetSuite Conduct data migrations from the current platform onto NetSuite. Collect requirements from clients for customizations/automation/integration Develop, test, and implement customized solutions for the NetSuite platform including scripts, workflows, and other customizations Create custom scripts using Suite Script 2.1 Communicate efficiently with functional consultants, project managers, and end-users Create custom forms, fields, searches, reports Install required bundles and 3rd party solutions Develop integrations with 3rd party systems with their APIs Conduct Testing Basic Qualifications: Previous experience working with NetSuite ERP and OneWorld is mandatory Experience creating custom scripts with Suite Script 2.0 & 2.1 2 years experience in JavaScript and other web programming languages 1-2 years of experience with Celigo Smart Connectors and Integrator.io Experience with configuration, sdf deployment and netsuite development Experience with advance pdf templates like check templates. NetSuite Suite Cloud Developer Certification is a plus Desired Qualifications: Experience in HTML, CSS3, JavaScript, jQuery, Backbone, Angular, AJAX, Suite Script Understanding of object-oriented concepts, abstraction/inheritance, as well as experience with object-oriented languages Data management preferred (SQL, XML, JSON) Web services preferred (REST, SOAP) How You’ll Embody Our Core Values At Plative, our core values shape how we work, collaborate, and grow. As part of the team, you will: Put Pe
TextNow is on a mission to make communications affordable and accessible for everyone. As a full MVNO operating our own mobile core network over LTE and 5G NSA, we have the unique advantage of controlling our network infrastructure end-to-end. We operate the HSS, PGW, and other critical network functions, giving us the flexibility to innovate and deliver exceptional service to millions of users. About the Role Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. T extNow is looking for a new SecOps team member to secure, monitor , and enable automated response within our infrastructure. What You’ll Do Ensure Secure & Reliable Systems: Design, implement, and maintain security-focused infrastructure to protect TextNow’s services while ensuring reliability and scalability. Security Automation & Infrastructure as Code: Develop and enforce best practices using Terraform, Ansible, Crowdstrike , and AWS security tools , ensuring secure configurations, automated compliance checks, and infrastructure as code. Threat Detection & Incident Response: Participate in an on-call rotation to respond to security incidents, investigate vulnerabilities, and implement proactive measures to prevent future threats. Work closely with engineering teams to remediate security risks. Monitoring & Logging for Security: Improve observability by implementing security monitoring solutions, logging best practices, and alerting mechanisms to detect anomalies and suspicious activity. Access Control & Identity Management: Manage IAM roles, permissions, and policies to ensure least privilege access and enforce security controls across cloud and internal systems. Collaboration & Security Advocacy: Wo
As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo
As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
At Scale, our mission is to develop reliable AI systems for the world's most important decisions. For 10 years, Scale has provided the high-quality data and full-stack technologies that power the world's leading models, and has helped enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. Public Sector engineers build the core product including the systems required to ingest and process federal datasets that support real-time decision-making in contested environments. As a New Grad Software Engineer on this team, you will own meaningful, mission-facing work from day one: shipping features, sitting with the government stakeholders who use them, and iterating fast. Example Projects Build multi-layered guardrails that keep agents safe and predictable in high-stakes federal environments Optimize data retrieval for agents, including RAG pipelines over large, heterogeneous federal datasets Build orchestration for fleets of asynchronous agents running long-horizon tasks Develop systems that automatically alert users to deviations and anomalies in incoming data Create interfaces that illustrate how an agent reached a decision, so operators can audit and trust its output Develop data pipelines and ML infrastructure that make previously siloed government data sources accessible to agents Build evaluation infrastructure that measures model reliability against mission requirements Ship full-stack tooling that lets analysts query, visualize, and explore mission data Deploy and harden applications into secure, air-gapped, and cloud-native government environments Requirements A graduation date in Fall 2026 or Spring 2027 with a Bachelor's degree (or equivalent) in a relevant field (Computer Science, EECS, Computer Engineering, Statistics) Product engineering expe
About Hexnode Hexnode, the Enterprise software division of Mitsogo Inc., was founded with a mission to simplify the way people work. Operating in over 100 countries, Hexnode UEM empowers organizations in diverse sectors. Fueling the transformation to a seamless ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Role Overview Drive high-velocity pipeline and revenue growth through a strategic partner ecosystem across the Philippines. You will own the full partner lifecycle: recruiting, enabling, and co-selling with top-tier Value-Added Resellers (VARs), Managed Service Providers (MSPs), Systems Integrators (SIs), and Distributors to win mid-market and enterprise deals . You will target key decision-makers in IT, Cybersecurity, and Endpoint Management . Responsibilities Identify, recruit, and onboard high-performing regional partners across Metro Manila, Cebu, Davao, and key IT-BPO hubs . Execute quarterly business plans ("Top 10" framework), joint pipeline generation campaigns, and regular Business Reviews (QBRs) . Enable partners with product demos, competitive battlecards against incumbents (e.g., Intune, Jamf, Ivanti, Workspace ONE), and pitch coaching . Assist partners in co-selling and closing multi-stakeholder enterprise deals (1,000 to 10,000+ endpoints) across key Philippine sectors (BPO/BFSI, Healthcare, Retail, Government) . Collaborate on partner incentives, co-marketing initiatives (MDF), and localized resale/referral structures . Manage partner deal registrations and maintain accurate sales forecasting to ensure consistent quota delivery . Act as an internal partner champion to align Solutions Engineering, pricing approval, and technical support resources . Required Experience & Expertise 5 to 10 years of B2B partner/channel sales experience in SaaS, Cybersecurity, Cloud, or IT Infrastructure . Proven track record of carrying a dedicated partner quota and managing 2-tier distribution models (
Integration meets innovation Celigo is the modern integration platform as a service (iPaaS) that simplifies how companies integrate, automate, and optimize processes. The platform has been ranked #1 in iPaaS user feedback by G2 for five quarters in a row. Purpose-built for mission-critical processes, Celigo offers unique tools such as runtime AI and prebuilt integrations tailored to resolve the biggest integration challenges, making Celigo incomparably easier to maintain. Celigo is looking for a results-driven Account Executive to own the full sales cycle, from prospecting through close, driving new business in the commercial segment, across Australia and New Zealand. The successful candidate will target and partner with fast-growing companies that will rely on Celigo to streamline operational excellence by providing a market-leading Integration Platform-as-a-Service (iPaaS) connecting cloud and on-premises software applications, allowing businesses to automate workflows and sync data across systems without heavy custom coding. In addition to understanding the demands of greater connectivity, the right candidate should bring a forward-thinking, AI-informed approach to how you work, compete, and win. In summary, if you have the mindset of thinking outside the box, positioning an end to legacy and tools with intelligent automation (to help consumers deliver significantly more without adding complexity to their business), we need to talk to you! What would you do if hired? Prospect and qualify inbound and outbound opportunities within an assigned territory; build pipeline through both direct outreach and Celigo’s partner ecosystem Conduct discovery to understand customer business challenges, integration requirements, and automation opportunities Position Celigo’s iPaaS platform as a strategic enabler for automation, scalability, and operational efficiency — including AI-powered capabilities like runtime AI and agentic workflows Work effectively and efficie
Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Summary: As a Principal Software Engineer at Appian, you will be the primary technical strategist, responsible for shaping the architectural foundation of our platform. Your role involves anticipating future challenges and implementing innovative solutions today. You will drive complex cross-functional initiatives, ensuring Appian remains a leader in the low-code and automation industry. Responsibilities: Define the long-term architectural vision and governance for the Appian platform. Identify systemic technical risks and lead task forces to resolve architectural bottlenecks. Develop internal tools and SDKs to simplify infrastructure and enhance developer productivity. Conduct deep-dive troubleshooting for complex production issues. Promote AI-native engineering practices and integrate AI features into the platform. Lead the development of core platform libraries and high-risk prototypes. Participate in the Architectural Guild and review high-impact design documents. Mentor lead and senior engineers, and represent Appian in the tech community. Ensure operational resilience with self-healing and highly available systems. Required Qualifications: Bachelor’s or Master’s degree in Computer Science, Information Technology, or related field. 15+ years of software engineering experience, with significant experience in architecting large-scale distributed systems. Strong understanding of data structures, algorithms, and design patterns. Proven transformational leadership in technological migrations or strategies. Expertise in Java and Cloud-Native ecosyst
Role Purpose We’re looking for a Staff/Senior Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Experience: 5+ years of industry experience in Machine Learning, with at least 3 years dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native: Hands-on ex
Machine Learning Engineer We’re looking for a Machine Learning Engineer with deep expertise in computer vision and biometrics to lead the design and scaling of face recognition systems in production. You’ll build and train models, and own ML systems end-to-end on AWS. The final job level for this role will be determined following the interview process. What You’ll Do Lead the design and development of computer vision systems for biometrics (face attributes, detection, quality, and recognition) Rigorous fairness analysis and benchmarking of biometric models across various datasets and operating conditions. Architect, train, and optimize models using PyTorch, Tensorflow, and/or JAX Own and evolve end-to-end ML pipelines, from data ingestion to deployment. Design automated pipelines (Airflow) for data ingestion and cleaning. You will be responsible for curating balanced training sets and generating synthetic data to address both quality and diversity gaps. Production Engineering: Own the path to production. Optimize models for low-latency inference (quantization, distillation, TensorRT/ONNX) and manage deployment on AWS. Mentor ML engineers, conduct code/design reviews, and drive technical best practices across the Computer Vision team. What We’re Looking For Experience: Multiple years of industry experience in Machine Learning, with at least 3 years dedicated to Biometrics or Face Analysis. Deep expertise in computer vision and biometrics, especially face recognition. Fairness & Ethics: You understand the sources of algorithmic bias in Computer Vision and have practical experience measuring and mitigating disparate impact. Strong Engineering: Expert proficiency in Python (both machine learning and vision libraries such as Pillow, OpenCV, PyTorch, etc). You write clean, modular, production-ready code. Systems Architecture: Experience designing end-to-end ML pipelines (Data to Train to Deploy) and working with workflow orchestrators like Airflow. Cloud Native:
Get new cloud operations system administrator jobs by email
Daily job updates · Unsubscribe anytime