We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing Identify and configure key metrics to detect incidents and quantify service health, availability, and performance Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Mentor early-career SREs and contribute to the team’s operational practices as it grows Qualifications Strong background in software development and operating distributed systems 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues Expertise in cloud infrastructure platforms, in
Jobiba hiring network
Linux System Administrator Jobs
742 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current linux system administrator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The database market is massive and MongoDB is at the head of its disruption. The MongoDB community is transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the forefront of innovation and creativity. Building on the rapid success and adoption of MongoDB, we are delivering applications and services that make it much easier to manage and scale database deployments. Our Cloud Technical Services Engineers provide their exceptional analytical skills and customer service to ensure that MongoDB users are successful with our suite of MongoDB application-layer products. We’re looking for individuals who want to dig into the details of how “big data” and “web-scale” systems are successfully assembled and operated every day by organisations of every size and flavour. Cloud Technical Services Engineers are a key part of our elite Technical Services organisation that supports MongoDB customers worldwide, with teams in locations including New York, Sydney, Gurugram, and Dublin. If you’re passionate about the opportunity to get comfortable working with the cloud every single day and be part of a team that works at the frontier of SaaS services and database systems, this is the role for you. We are looking to speak to candidates who are based in Sydney for our hybrid working model. Responsibilities As a Cloud Technical Services Engineer, you’ll be advising customers on strategies and documented practices for making best use of our cloud products. This will require you to translate technical concepts and patterns into generalist terms for our customers, helping them understand, install, and use those applications well. You’ll also troubleshoot technical problems with these products, and be an advocate for our users’ needs, collaborating with the MongoDB product management and development teams on their behalf.&nbs
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Technology, Data and Intelligence Team Message Okta’s Technology, Data and Intelligence (TDI) team delivers the systems, tools, and services that power internal operations across the company. From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will help build, improve, and maintain our cloud platform services by designing and implementing complex cloud-based engineering enablement systems. With a strong focus on automation, testing, and operational excellence, you will deliver foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale. What you'll be doing Secure Cloud Infrastructure & Pipelines: Design, build, and modernize scalable cloud environments and development tools while strictly enforcing security policies and standards for regulated environments. Cross-Functional Collaboration & Advocacy: Partner with software engineering teams to champion DevOps and SRE best practices, deliver excellent internal customer service, and actively contribute to Agile workflows (e.g., demos, architecture sessions). Technical Documentation & Operations: Create and maintain comprehensive technical documentation, including network diagrams, runbooks, and disaster recovery procedures to en
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, our motto is "Always On" and nowhere do we embrace that more than in Technical Operations. We strive to build the most reliable and performant systems on the planet through the skillful use of automation. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. You will work on: ● Mentoring, managing, and leading a team of SRE’s with a broad range of expertise and experience. ● Being an evangelist and advocate for security best practices, leading initiatives and projects to strengthen our security posture for our most critical infrastructure. ● Responding to production incidents, driving us to remediation as quickly as possible and determining how we can prevent them in the future. ● Triaging and troubleshooting complex production issues to ensure reliability and performance. ● Working closely with our stakeholders across the organization to ensure our new capabilities are aligned to our competing constraints of reliability, security, and delivery velocity. ● Partnering directly with recruiting and people ops to hire and retain the best talent in the world. ● Keep sharp eyes on our metrics, including vulnerability scanning and security posture, 
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Data Platform team builds and operates the systems that centralize all of Coinbase's internal and third-party data, making it easy for teams across the company to access, process, and transform that data for analytics, machine learning, and powering end-user experiences. As an engineer on the team you'll contribute across the full spectrum of our systems, from foundational processing and storage to scalable pipelines to the frameworks, tools, and internal applications that make data easily and efficiently available. We have a number of open roles within the Data Platform team, each with a different focus area; read on to see if you'll be a good fit as a Data Engineering Platform Engineer. This role sits between our platform and the analysts, data scientists, and data engineers who use it every day. You'll build the SDKs, self-service apps, and abstractions that let thousands of pipelines and analyses run without those users needing to understand the plumbing underneath. What you'll do: Build clean software abstractions and modular, reusable components that let us support thousands of pipelines and beyond as we scale. Build and maintain data integration and processing SDKs used by data teams and internal services across Coinbase. Design and build self-service applications that let analysts and data scientists create, manage, and troubleshoot their own pipelines wi
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We’re seeking an exceptional Principal-level Offensive Security Engineer focused on deep, hands-on penetration testing of OpenAI’s agent-powered products, infrastructure, and model-integrated application surfaces. You’ll assess complex systems end to end, identify realistic vulnerabilities, validate exploitability and impact, and partner closely with engineering teams to drive durable fixes. This role will be primarily focused on continuously testing our agent-powered products like Codex and Operator. These systems are uniquely valuable targets because they’re rapidly evolving, can perform sensitive actions on behalf of users, and have large, diverse attack surfaces. You will play a crucial role in securing our agents by finding vulnerabilities that emerge from the interactions between the applications, infrastructure, tools, and models that power them. You’ll have the chance to not only find vulnerabilities, but actively drive their resolution, build reusable testing approaches, automate offensive security workflows with cutting-edge technologies, and use your attacker perspective to improve the security of OpenAI’s products. In this role you will: Conduct deep penetration tests of OpenAI’s agent-powered products, including web applications, APIs, cloud services, identity and authorization flows, CI/CD systems, and model-integrated product surfaces. Continuously hunt for exploitable vulnerabilities in the interactions between the appli
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity The Forward Deployed Engineering (FDE) team tackles some of Postman's most strategic technical challenges. We partner directly with enterprise customers to solve problems that don't fit neatly into existing product boundaries, building production-grade solutions that often become core product capabilities. As a Sr. Forward Deployed Engineer, you'll operate at the intersection of engineering, product, and customer success. You'll deploy Postman's critical infrastructure into customer environments, solve complex distributed systems challenges, and build the enterprise foundations that enable customers to safely adopt AI at scale. This isn't consulting or professional services. You're an engineer building production systems alongside customers, turning real-world deployments into durable platform capabilities used by thousands of organizations. This team will build 0-to-1 products from scratch - innovative, strategic initiatives designed to unlock hundreds of millions of dollars in new revenue. If you enjoy difficult engineering problems, working directly with customers, and building products from the field
As a Software Solution Architect, NVIS at NVIDIA, you will lead the transformation of AI infrastructure. Our NVIS team focuses on developing the next generation of NVIS Central, an agentic software platform with tools, services, and AI agents that automate, simplify, and speed up the work of our delivery organization. This role offers an outstanding chance to create and build LLM-powered agents that improve execution visibility, cut down manual tasks, and standardize workflows. These efforts allow NVIS to grow quickly and with high quality. Join us to bring up, validate, optimize, and upgrade large-scale AI Factory infrastructure for some of the world’s most advanced accelerated computing environments! What you'll be doing: Compose, build, and productionize agentic AI solutions, tools, and applications for the NVIS delivery organization. Develop LLM-based agents, skills, tool-calling workflows, orchestration logic, backend services, APIs, data pipelines, and automation features as part of NVIS Central. Translate field, delivery, operations, and product needs into clear technical builds, agent workflows, and working software. Develop agents that can reason across project data, knowledge bases, operational systems, logs, reports, and delivery workflows. Build workflows that help NVIS teams identify risks, summarize project status, automate repetitive tasks, improve readiness visibility, and simplify handoffs. Work with timely engineering, retrieval-augmented generation, context management, agent memory, function calling, evaluations, and guardrails to build reliable AI systems. Integrate LLMs and agents with internal systems, project data sources, knowledge repositories, reporting tools, and operational workflows. Collaborate closely with software developers, architects, product managers, DevOps/SRE, and NVIS field teams to
Are you looking for an opportunity to help solve one of today's biggest business challenges? AI is changing the pace of business, and organizations everywhere are struggling to help their workforce, partners, and customers keep up. At Litmos, we're building the Learning Acceleration Platform that helps organizations build human capability faster—and we're looking for people who are passionate about making a meaningful impact for customers while growing alongside a collaborative, people-first team. Litmos is the Learning Acceleration Platform that helps organizations build capability faster, adapt at the speed business changes, and scale learning to anyone, anywhere. Combining an intuitive platform, AI-powered capabilities, trusted content, expert services, and a broad ecosystem of integrations, Litmos helps organizations accelerate workforce productivity, improve customer adoption and retention, enable high-performing partners, and reduce organizational risk through continuous learning. Organizations such as Hewlett Packard Enterprise, Graco, Sabre, and Russell Mineral Equipment trust Litmos to accelerate learning across their workforce, partners, and customers. Today, more than 11K customers with 30 million learners across 150 countries and 37 languages use Litmos to build the capabilities their organizations need to succeed. Backed by Francisco Partners, one of the world's leading technology investment firms, we're investing in the future of learning—and the people who are building it. Learn more at www.litmos.com . Network Operations Engineer About the Role We're looking for a Network Operations Engineer to join our CloudOps team and keep our infrastructure secure, patched, and available around the clock. You'll monitor and support a multi-cloud environment, primarily Azure, with an AWS footprint, spanning Windows, Linux, and SQL Server systems. Coverage runs follow-the-sun, so you'll be part of a rotation that ensures alerts never go u
Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! Job Description - Technical Support Engineer (Rotational Split Shift: Sat - Wed) We are looking for a Technical Support Engineer to assist our rapidly growing customer base. As part of our Technical Support team you will be the primary point of contact for Couchbase customers for all technical issues. NoSQL databases are the answer to the demand for high speed, highly available, extremely dynamic data storage and Couchbase is at the forefront of this technology. Working in Couchbase Technical Support you’ll acquire highly coveted skills, essential to this technology, by navigating the world of NoSQL databases. This includes working with Golang, learning about the principals and concepts of distributed systems, understanding what goes into good NoSQL database design, mobile data convergence, and becoming proficie
SIXGEN’s mission is to deliver agile, mission-ready cybersecurity solutions that empower government and critical infrastructure organizations to stay ahead of advanced cyber threats. We combine innovation, deep expertise, and cutting-edge capabilities to uncover vulnerabilities, protect vital systems, and ensure operational superiority in an ever-evolving digital landscape. POSITION OVERVIEW Position: Software Engineer Job Type: Full-time Location: Ft. Meade, MD Clearance Requirements: Active TS/SCI with Polygraph required Work Schedule: 5 days per week onsite Travel Requirements: None Experience: 14+ years WHAT YOU'LL DO We are seeking a Software Engineer to support the implementation, troubleshooting, maintenance, and ongoing support of agile software development efforts focused on modernizing systems. This role requires applying advanced software engineering principles, theories, and concepts while contributing to the development of new approaches and solutions. The ideal candidate will tackle highly complex technical challenges, operate with significant autonomy, and provide technical leadership and mentorship to team members when needed. KEY RESPONSIBILITIES Software Development Design, develop, modify, and implement software applications using agile development methodologies. Write source code for new applications and enhance existing software solutions. Utilize both front-end and back-end technologies to develop complete software solutions. Support the modernization, maintenance, and enhancement of existing systems. Engineering & Technical Problem Solving Apply advanced software engineering principles to solve complex technical challenges. Develop innovative approaches and solutions to meet mission requirements. Contribute to software architecture, integration efforts, and application development activities. Collaborate across cross-functional teams to deliver effective solutions. Team Collaboration & Leadership Work closely with technical teams in a
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE The CPBU Manageability team is responsible for Middleware of FlashArray and FlashBlade products. We are building a scale-out all-flash file and object store, designed for the modern world. To really understand how our customers work with data, we are deeply immersed in AI, modern backup, log analytics with Splunk and Elastic, data pipeline with Kafka, cluster computing with Spark, and many more use cases. You will love it on the CPBU Manageability team if you: want to understand how modern applications work with data and how we can make it better. are ready to dive into a complex problem and be the one who will drive it to a resolution. enjoy working with distributed systems, algorithms, operating systems, Linux kernel, database internals, hypervisors, containers, compilers and hardware… or at least some of those. want to work with other great engineers and develop or refine skills that will serve your entire career. enjoy working in a collaborative team environment in an open office. If this describes you, let's talk! You can take a part in changing how the world works with data. WHAT YOU'LL DO Own and deliver innovation end-to-end, from concept to shipped product Design, develop and maintain customer-facing and internal-facing API and command line interface for end to end configuration and management of FlashArray and FlashBlade products using Java, Python and beyond Experimenting with new technologies and archi
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE The FlashArray team builds an innovative, high-performance, and highly available portfolio of products designed for demanding, mission-critical applications. While we deliver a hardware storage array, over 90% of our engineering team focuses on software engineering. Our customers value FlashArray for its simplicity of management, continuous feature upgrades, and ability to stay on the cutting edge without downtime. Recently, we extended these capabilities into the public cloud with CloudSnap and Cloud Block Store for AWS, enabling customers to leverage cloud agility for both traditional IT and cloud-native applications. WHAT YOU’LL DO Design & Implement: Create innovative algorithms and technologies for high-performance systems targeting six-nines (99.9999%) reliability. End-to-End Ownership: Lead feature innovation from initial concept through to shipped product. Problem Solving: Analyze and resolve complex technical challenges through persistent problem-solving and technical insight. Collaborate & Deliver: Partner closely with smart, collaborative peers to deliver features that directly enhance customer experience and satisfaction. Learn & Grow: Continuously expand domain expertise in systems software within a supportive, knowledge-sharing environment. WHAT YOU BRING Software Development Experience: 3+ years of professional development experience using C, C++, Python, Go, Java, or related prog
Get new linux system administrator jobs by email
Daily job updates · Unsubscribe anytime