About the team The Agent Enablement AI Deployment Engineering (ADE) team works across engineering, product, design, partnerships, and strategic customers to grow an open ecosystem of agent-enabled sites and services. We help partners adopt the OpenAI tech stack related to identity, permissioning, agent-auth primitives so users can safely connect ChatGPT and Codex to the tools, services, and workflows they already use. Our team also works with external partners on defining the standards for agent access, marketplace offerings as well as other agent enablement initiatives to ensure users of ChatGPT and Codex go from intent to task completion seamlessly. About the role We are looking for an AI Deployment Engineer to help strategic partners design, build, validate, launch, and operate agent enablement integrations across web applications, connectors, APIs, CLIs, MCP servers, and developer tools. This is a hands-on, partner-facing product engineering role for someone who can contribute to the platform itself, lead sophisticated technical engagements, and turn ambiguous identity and agent-workflow requirements into secure, production-ready integrations. You will work across partner product and engineering teams and OpenAI’s product, engineering, design, partnerships, legal, policy, security, support, and go-to-market teams. You will identify high-value user journeys, choose the right integration path, prototype and review architectures, write code, run evaluations and dogfood, trace failures end to end, guide launch and rollout, and support post-launch iteration. The best person for this role moves fluidly between full-stack code, OAuth/OIDC and identity systems, product judgment, project leadership, and clear communication with engineers and executives. This role is a fit for a product-minded engineer who wants to stay close to users and partners while going deep on authentication, permissions, reliability, safety, and developer experience. The principle objective is to
Jobiba hiring network
Servers Jobs
846 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current servers jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Overview At Okta, we're building a world where anyone can safely use any technology, and we're specifically focused on protecting the future of work. We bring simple and secure access to people and organizations everywhere, promising to not only protect the identities of our customers' workforce and users but to ask, “what more can we make possible?” As a Staff Product Designer for Okta, you’ll work with a globally distributed design team to apply your craft and expertise on a range of products and services. Our culture provides the opportunity to explore the technologies we’re researching and investing in. Some of your work will be functional, practical user experience design, where you will help our design team drive cohesive user experience across Okta’s portfolio of products. Other work will be speculative, generative explorations that you carry out in collaboration and partnership across the organization to flesh out, frame, and help prioritize solutions in an ever-evolving and complex, often ambiguous, problem space. The Finer Print You will lead user experience strategy and execution for our Okta Research & Development (ORD) business unit. Specifically, ORD will solve every identity use case at work – for all types of people (employees, business partners, contractors, frontline workers), for all types of resources (SaaS, servers, privileged resources, machines, etc.), for the largest and most secure businesses in the world.
Spaulding Ridge is an advisory and IT implementation firm. We help global organizations get financial clarity into the complex, daily sales, and operational decisions that impact profitable revenue generations, efficient operational performance, and reliable financial management. At Spaulding Ridge, we believe all business is personal. Core to our values is our relationships with our clients, our business partners, our team, and the global community. Our employees dedicate their time to helping our clients transform their business, from strategy through implementation and business transformation. What You Will Do and Learn: The Cybersecurity Operations Analyst will support the Cybersecurity Manager in maintaining and improving Spaulding Ridge's security posture. This entry-level role is ideal for someone looking to develop a career in cybersecurity while gaining hands-on experience across security operations, vulnerability management, identity and access management, cloud security, and compliance. The analyst will assist in monitoring security events, responding to security alerts, supporting security projects, and maintaining security documentation while learning from senior members of the Technology & Security team. Security Operations Monitor and analyze security alerts from SIEM, EDR, Firewall, WAF, IDS/IPS, email security, and cloud security platforms. Conduct and Support threat hunting and detection of tuning activities. Monitors open-source and proprietary threat intelligence feeds daily to research threat actor tactics, techniques, and procedures (TTPs). Support the investigation and triage of security alerts and incidents under the guidance of the Cybersecurity Manager. Assist with incident response activities, including documentation, evidence collection, and follow-up actions. Help maintain security playbooks and operational procedures. Vulnerability Management Assist with vulnerability scanning across endpoints, servers, and cloud environments.
Senior Infrastructure Architect — Enterprise Observability and Automation Description - Job Summary Senior individual contributor responsible for the architecture, implementation, and operational ownership of enterprise observability, monitoring, and automation platforms across HP's global IT environment. This role modernizes infrastructure visibility capabilities while ensuring operational stability, security, and compliance. Serves as a technical and operational bridge between infrastructure engineering, cybersecurity, SOX/compliance stakeholders, automation teams, and external technology partners — leading complex initiatives such as platform migrations, enterprise integrations, and governance enablement. Responsibilities Enterprise Observability and Monitoring Application owner and senior technical authority for enterprise monitoring and logging platforms (Datadog, Splunk), including platform governance, roadmap alignment, and operational oversight. Lead enterprise-scale monitoring platform migrations, including architecture design, agent strategy, data ingestion models, vendor coordination, and deployment across 5,000+ servers. Define standards for alerting, dashboards, observability data quality, and integration with ITSM platforms (ServiceNow). Design and manage multi-org Datadog architecture, including org structure, RBAC, SSO/SAML, secrets management, and cybersecurity compliance. Oversee SNMP-based monitoring of storage and network devices, including device profiling, syslog/event integration, and NetFlow collection. SOX Compliance and IT Governance SOX control owner for enterprise monitoring applications — approve monthly reviews, participate in internal/external audits (EY), and maintain ITGC/SOX compliance. Provide audit evidence, walkthrough docu
Field Technical support Associate Description - Job Summary • This role is responsible for engaging with customers to enhance satisfaction by meeting their needs as well as delivering standardized services that meet established quality benchmarks. The role performs installations, maintenance, and repairs on customer equipment according to industry standards. The role ensures quality services, analyzes IT issues, maintains documentation, and provides site support. The role also acquires skills, adheres to guidelines, supports plans, and performs assigned tasks while maintaining confidentiality. Responsibilities • Engages in interactions with customers in adherence to established procedures, aiming to ascertain and enhance customer satisfaction by clarifying their needs and ensuring they are met. • Conducts installations, reinstallations, maintenance, and repairs on customer equipment, following industry standards and best practices. • Delivers standardized services that meet established quality benchmarks. • Leverages data from reliable sources to ensure precise alignment with customer product requirements. • Analyzes, troubleshoots, and resolves issues within IT infrastructure, including enterprise systems, servers, storage, and networking. • Maintains departmental documentation on work orders, software, inventory, and other paperwork required. • Provides site support for customer break fix activity and provides technical assistance to third-party and channel authorized service providers. • Acquires job skills and completes routine assignments/tasks, adhering to guidelines, and ensuring confidentiality in all dealings with company data. • Assists in implementing new processes, supports department-level operational plans, and shares technical information with colleagues and clients. • Solves defined problems using es
Field Technical Support Description - Job Summary • This role is responsible for engaging with customers to enhance satisfaction by meeting their needs as well as delivering standardized services that meet established quality benchmarks. The role performs installations, maintenance, and repairs on customer equipment according to industry standards. The role ensures quality services, analyzes IT issues, maintains documentation, and provides site support. The role also acquires skills, adheres to guidelines, supports plans, and performs assigned tasks while maintaining confidentiality. Responsibilities • Engages in interactions with customers in adherence to established procedures, aiming to ascertain and enhance customer satisfaction by clarifying their needs and ensuring they are met. • Conducts installations, reinstallations, maintenance, and repairs on customer equipment, following industry standards and best practices. • Delivers standardized services that meet established quality benchmarks. • Leverages data from reliable sources to ensure precise alignment with customer product requirements. • Analyzes, troubleshoots, and resolves issues within IT infrastructure, including enterprise systems, servers, storage, and networking. • Maintains departmental documentation on work orders, software, inventory, and other paperwork required. • Provides site support for customer break fix activity and provides technical assistance to third-party and channel authorized service providers. • Acquires job skills and completes routine assignments/tasks, adhering to guidelines, and ensuring confidentiality in all dealings with company data. • Assists in implementing new processes, supports department-level operational plans, and shares technical information with colleagues and clients. • Solves defined problems using established
Experienced Engineering Technical Specialist Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking Experienced Engineering Technical Specialists (Level 3) to join the team in Berkeley, MO. This position supports Risk Managed Framework compliance across multiple avionics test labs as well as the administration of embedded PCs and host PCs on avionics test benches. The Lab Infrastructure team is responsible for the design, modification, and troubleshooting of all test benches used in the avionics' lab to support systems integration testing on the F-15. Position Responsibilities Build and configure embedded PCs and host PCs on test benches Support RHEL 9 Installation, configuration, patching and maintenance of Linux server environments Support containerization Provide active directory management and file server management Troubleshoot user problems related to hardware, software, or processes Conduct monthly OS patching and anti-virus updates Create user accounts and reset passwords Run backups on workstations and servers Maintain compliance with Risk Managed Framework and NISPOM rules Keep facility signage up to date Maintain lab access list Travel may be required up to 10% of the time; domestically and/or internationally depending on business needs. This position requires an active U.S. Secret Security Clearance (U.S. Citizenship Required). (A U.S. Security Clearance that has been active in the past 24 months is considered active) Basic Qualifications (Required Skills/Experience) <ul
Do work that matters. At AlertMedia, we help organizations protect their people, operations, and brand. Our modern Risk Intelligence and Response platform empowers teams to detect emerging threats, assess impact, and respond with confidence. We believe building resilience should be simpler—and it starts with bringing critical information and workflows together in one unified platform. Our core values drive us in our important mission of keeping people safe & informed: We’re humans not robots Customers always come first We work better together Simplicity is our strength Our reputation is priceless Hard work pays of The IT Support Specialist at AlertMedia will work with other IT staff to provide our employees with fast and friendly support with all their technology needs. This includes diagnosing, repairing, maintaining and upgrading Windows & Mac computers, servers, phones and critical software applications as necessary to ensure smooth operation. Who you are: You are a highly motivated, organized team player who is passionate about technology, customer service, and operational follow-through. You bring exceptional attention to detail, quickly get to the “why” behind every request or issue, and proactively keep work moving through clear action and timely communication. You are a self-starter who enjoys a challenge, thrives in a fast-paced environment, and takes pride in providing efficient, thoughtful support to fellow employees. You love figuring out how computers work, learning new things, following the latest OS updates, and have a healthy respect for patch management. What you get to do every day: Provide first level support of end users technology related issues in a prompt, courteous manner Respond to incoming calls, tickets, and Slack messages regarding computer or equipment problems, and software issues Guide us
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides developer-first observability to over 4 million developers worldwide. The Events Analytics Platform (EAP) team is at the heart of that mission: it powers how all of Sentry's event data, such as errors, transactions, spans, profiles, replays, and metrics, is stored, queried, and analyzed. It also powers Sentry's latest AI push, Seer. The EAP team makes it possible for developers to efficiently search and debug across massive volumes of data, providing the context needed to understand and fix issues quickly. This team is also a cornerstone of Sentry's long-term strategy to become a context assembly and telemetry platform that unifies different signals so developers can see the complete picture. As an engineering manager on the EAP team, you will lead a group of engineers building and scaling one of Sentry's most critical data platforms. You will be responsible for driving architectural evolution, ensuring system stability, and mentoring a talented team. This is a highly visible leadership role with direct ties to Sentry's long-term product and platform strategy. What you'll do Grow and develop a team of engineers with high expectations for ownership and impact Set the technical and strategic direction for the team, balancing short-term stability with long-term architectural evolution Drive development of core EAP features, including support for complex analytical queries, dynamic routing logic across fidelity levels, storage and compute separation, and modern patterns for analytical storage Ensure EAP can support the workload demands of AI agents and MCP servers that unlock new AI capabilities fo
Citi is evolving its current commerce offerings and branching into new avenues of growth by offering merchants access to high value cardmembers across Citi’s digital properties and partner placements. By combining Citi’s scale, trusted customer relationships, and real purchase insights, the business delivers advertising that is relevant to customers and drives measurable, incremental impact for advertisers. If you're early in your career and have been working within AdTech, MarTech, or data environments and are ready to deepen your technical expertise while helping build a platform from the ground up, this is the role. The Junior Solution Architect will support the design and implementation of Citi Commerce’s AdTech and data infrastructure by translating defined architecture into scalable, working technical components. This is a highly hands-on, detail-oriented role focused on execution—ensuring platforms are properly integrated, configured, and optimized to support campaign activation, ad serving, and measurement. This role is ideal for an early-career technologist with AdTech or MarTech exposure who enjoys working across systems, troubleshooting integrations, and enabling seamless data flow across platforms. Responsibilities: Support implementation of Citi Commerce’s end-to-end AdTech and data architecture across platforms Assist with integrations across ad servers, DSPs, CDPs, and measurement platforms (e.g., GAM, DV360, LiveRamp) Configure and maintain tagging, tracking, and pixels to ensure accurate campaign execution and reporting Help build and manage data pipelines, ensuring reliable data ingestion and flow across systems Troubleshoot technical issues related to integrations, tracking discrepancies, and data inconsistencies Assist in QA and validation of platform integrations, tagging, and measurement frameworks Document system c
We build and operate the compute infrastructure our researchers run on, supporting large-scale processing of historical market data and model training on our own hardware across multiple data centers. Our environment includes bare-metal Linux, virtualization, storage, and GPU clusters, where performance, reliability, and predictable system behavior are critical. Our Infrastructure team covers monitoring and automation, distributed storage, hardware and OS provisioning, GPU clusters and workload scheduling, high-speed networking, L2/L3 Linux support, and security engineering. Engineers here own their tasks end to end, so there's room to go deeper in your area and pick up the parts you haven't touched yet. We’re looking for a Linux Infrastructure Engineer who can work hands-on with server and cluster environments, from deployment and configuration to performance tuning, troubleshooting, and ongoing improvement What You’ll Be Doing: Deploying, configuring, and maintaining Linux-based bare-metal servers across our data centers Building and operating clustered environments, including virtualization, storage, GPU compute, and database clusters Troubleshooting complex Linux, hardware, networking, and cluster-level issues Performance tuning for throughput, latency, stability, and resource utilization Monitoring infrastructure health and performance, identifying bottlenecks, and preventing recurring issues Supporting the full server lifecycle: provisioning, setup, upgrades, and maintenance Improving reliability and predictability during failures, maintenance, and scaling Automating provisioning, configuration, and operational tasks, primarily using Ansible and scripting What We Look For In You: Strong hands-on Linux administration and troubleshooting experience Production experience with on-premise, bare-metal infrastructure Good understanding of Linux performance and bottleneck analysis Experience with: infrastructure monitoring and troubleshooting production issues,
WHO ARE WE? We are a bunch of super enthusiastic, passionate, and highly driven people, working to achieve a common goal! We believe that work and the workplace should be joyful and always buzzing with energy! CloudSEK , one of India’s most trusted Cyber security product companies, is on a mission to build the world’s fastest and most reliable AI technology that identifies and resolves digital threats in real-time. The central proposition is leveraging Artificial Intelligence and Machine Learning to create a quick and reliable analysis and alert system that provides rapid detection across multiple internet sources, precise threat analysis, and prompt resolution with minimal human intervention. Founded in 2015, headquartered at Singapore, we are proud to say that we’ve grown at a frenetic pace and have been able to achieve some accolades along the way, including: CloudSEK’s Product Suite: CloudSEK XVigil constantly maps a customer’s digital assets, identifies threats and enriches them with cyber intelligence, and then provides workflows to manage and remediate all identified threats including takedown support. A powerful Attack Surface Monitoring tool that gives visibility and intelligence on customers’ attack surfaces. CloudSEK's BeVigil uses a combination of Mobile, Web, Network and Encryption Scanners to map and protect known and unknown assets. CloudSEK’s Contextual AI SVigil identifies software supply chain risks by monitoring Software, Cloud Services, and third-party dependencies. CloudSEK’s AIVigil is an AI-native Attack Surface Monitoring platform that continuously discovers, monitors, and secures exposed AI infrastructure, MCP servers, leaked AI credentials, vector databases, agentic workflows, and shadow AI across the internet. Key Milestones: 2016 : Launched our first product. 2018 : Secured Pre-series A funding. 2019 : Expanded operations to India, Southeast Asia, and the Americas. 2020 : Won the NASSCOM-DSCI Excellence Award for Security Product Company
About Backblaze Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands. Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $100m in revenue and is the leading specialized storage cloud - managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals. But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a Sr. Software Engineer to join our team! What You'll Do: You will work on our storage platform which supports the B2 Object Storage and Computer Backup products, responsible for reliably storing the customer data. You will help improve durability, increase throughput, and drive architectural changes to support the long term requirements of our service. The Right Fit: 7+ years of OOP (Java, Rust, C++, C#, etc) in an enterprise environment (personal project work not included) Experience building large scale software systems running on thousands of servers Comfortable with all aspects of the software development lifecycle, including design, implementation, testing, and rollout Exhibits curiosity and seeks to understand before offering solutions Applies analytical thinking to ambiguous problems Bonus points for: Proficiency in Java or Rust Experience building and interpreting models, dashboards, or reports Familiarity with hard drive performance Understanding of Linux internals and file systems At this point, we hope you're feeling
Get new servers jobs by email
Daily job updates · Unsubscribe anytime