Jobiba hiring network

Senior Infrastructure Engineer Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

E
19 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Our Senior Applied AI Engineer builds and operate production-grade AI systems that extract meaning from large-scale unstructured document collections, enabling enterprise data discovery classification, and governance. This role owns the full lifecycle of graph intelligence solutions — from problem definition and data modelling, to building and enriching knowledge graphs, and deploying ML- and LLM-assisted analytics in production. The focus is on semantic and contextual analysis of unstructured data to uncover relationships, patterns, and insights that support AI safety, security, and compliance requirements. WHAT YOU'LL DO Design, build, and deploy graph-based AI solutions, combining knowledge graphs , LLMs, and ML models applied to large-scale unstructured data Define and own data pipelines that extract, transform, and enrich entity relationships into production-grade knowledge graphs Integrate LLMs and ML models into text processing pipelines for classification, embedding generation, document similarity, and semantic analysis Design, deploy, and operate graph and vector databases to support retrieval, reasoning, and analytics Optimize models and inference pipelines for production constraints including latency, throughput, cost, and infrastructure Deploy, monitor, and iterate on ML systems in production environments ensuring reliability and continuous integration Drive architectural decisions and tech

pythonawsdocker
View job →

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . As a Software Dev Senior Engineer , you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error budgets, and continuously reducing toil through engineering. Key Responsibilities: Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs ) for all critical services, partnering with product and engineering leadership. Own the error budget framework: track consumption, enforce error budget policies, and drive reliability investments when budgets are at risk. Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into pro

pythonsqlpostgresql
View job →

NVIDIA is well positioned as the 'AI Computing Company', our GPUs being the brains that power modern Deep Learning software frameworks, accelerated analytics, modern data centers, and driving autonomous vehicles. We are looking for a Senior Software QA Test Development Engineer to join in the mission of crafting a distributed technology for all NVIDIA teams that remotely manage 10s of 1000s of resources in a simple and controlled fashion, allowing engineers to focus on engineering and automation, rather than being burdened by manual operational tasks. SWQA test developer engineers at NVIDIA are responsible for creating test plans, execution, and reporting, as well as developing scripts for test automation, designing and developing tools for the QA team, and developing integration tests for validation. As a test developer, you must identify weak spots and constantly design better and more creative test plans to break software and identify potential issues. You will have a huge impact on the quality of NVIDIA's products. The ideal candidate must have strong programming skills and hands-on experience using AI development tools to improve quality and productivity across the end-to-end QA workflow. This includes leveraging AI assistants for test automation, code generation, debugging, and enhancing testing efficiency. During the interview process, we will assess your ability to effectively use AI development tools and evaluate your programming capabilities to ensure you can deliver high-quality solutions. What you’ll be doing: Architect, implement, and evolve scalable agentic end to end SWQA workflow, automated test frameworks, infrastructure, and tooling for complex software products. Define test strategy and quality gates across functional, integration, regression, reliability, and release-validation workflows. Build and maintain high-value automated coverage for Linux-based, co

dockerkuberneteslinux
View job →

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our Team Do you want to be an Information Security Lead at GoDaddy? GoDaddy’s Security organization is looking for a Cloud Security Engineer. We work out large-scale and cross-company security challenges while ensuring that partnership with the development and operational communities remains front of mind. At GoDaddy, Security Engineers apply their strong hands-on technical skills to craft scalable solutions for multiple problems. You must communicate with GoDaddy Engineering teams, perform security assessments, prioritize security risks, and design. We, as a team, implement high-quality security engineering solutions! What you'll get to do... The Senior Cloud Network Security Engineer will play a crucial role in designing, building, and securing large-scale, distributed cloud environments that support GoDaddy Services. This role operates at the intersection of cloud infrastructure, security architecture, and engineering execution. The successful candidate will collaborate closely with service teams, security leaders, and compliance partners to embed security-by-design principles into cloud services and internal platforms. This position requires advanced technical expertise in cloud-native security controls. It also needs a strong understanding of threat models in hyperscale environments. Additionally, it involves influencing architecture decisions across multiple teams. The role demands hands-on engineering, good judgment in ambiguous situations, and proficiency at translating security requirements into scalable, automated solutions. Bui

pythonjavaaws
View job →
V
Vanta
📍 United States• Full-time
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Corporate Engineering team is the infrastructure layer that keeps 1,500+ people connected, secure, and moving fast. AI is now central to how that happens, and this role leads the team that owns the backbone underneath it. Corporate Engineering owns the shared AI platform for Vanta's internal systems: identity and access for AI tools, cost visibility and controls, sanctioned tooling and guardrails, and the enablement that helps people use those tools well. We do not own product AI, which stays with Engineering, and we do not replace the AI work happening inside individual departments. We build the engineering layer that makes all of it safer, cheaper, and better supported. The company has moved fast. AI assistant use is widespread across every function, MCP infrastructure is live and self-serve company-wide, internal apps ship on a hosted platform, and AI spend has grown to the point where attribution and guardrails genuinely matter. What does not exist yet is a single owner for that platform layer. That is this role. As Sr. Manager, AI Engineering, you will lead a small, senior team, write the charter for what Corporate Engineering owns versus enables, and build the platform that lets the rest of Vanta adopt AI quickly without accumulating cost, risk, or duplication. What you’ll do as a Senior Manager, AI Corporate Engineering at Vanta: Lead, coach, and grow a senior team spanning platform engineering and technical program management. Set clear direction, hold a high bar, and give the team explicit permission to push back with data. Own the AI access layer end to end: identity and authentication for AI tools, connector

ci/cdrestai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role Our robotics team is growing and we are looking for a Software Engineer to join our Sensor Data and Calibration team. We are searching for an engineer with robotics and machine learning expertise to develop synthetic sensor simulation models and algorithms. The ideal candidate has hands-on experience in the research, development, and implementation of machine learning methods (e.g., NeRF or Gaussian splatting) for generating synthetic sensor data (photorealistic images, realistic lidar and/or radar, etc.). About the Work Research, develop, and implement state-of-the-art synthetic sensor simulation methods Analyze and characterize the realism and utility of synthetic sensor data Answer critical questions about sensor data and autonomy performance Collaborate with stakeholders across autonomy, infrastructure, and systems teams on map needs and requirements Role is scoped as a Senior/Staff IC with the flexibility to grow into

pythonmachine learningai
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform. You Will: Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with

pythonsqlaws
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform. You Will: Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with

pythonsqlaws
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Boston, New York City, Raleigh, Miami, Pittsburgh or remotely in the United States while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure g

pythonmongodbaws
View job →

The Team This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design & build complex systems, operate with autonomy and act as owner for everything you do. The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers. The ideal candidate should Have 5+ years of experience running critical systems at scale Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”) Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment A strong understanding of how to run a large scale Linux environment, including low level fundamentals Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python) Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc) Special Requirements: Be a US Citizen Expectations Participate in the development of a reliable and resilient multi-cloud platform that hosts business critic

pythonmongodbaws
View job →

The Team MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of either our Dublin or Cork office or remotely in Ireland. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Build for reliability, making services and infrastructure avail

pythonmongodbaws
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Toronto or Montreal office or remotely in the Canada while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineerin

pythonmongodbaws
View job →
O
Okta
📍 London• Full-time• From £163K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. As a Senior Manager, Solution Engineering, you will: Hire, coach, and inspire a high-performing team of Solution Engineers, driving career development, performance management, and continuous skill enablement. Support and coach your SEs as they partner with Account Executives to discover customer pain points, develop opportunities, and secure the technical win. Partner directly with senior customer technical leaders on strategic opportunities to articulate Okta’s platform value and secure technical buying decisions. Build deep relationships with key partner technical sales leaders to establish Okta as their preferred identity security standard. Partner with aligned Area Sales Directors to design, execute, and scale Go-To-Market strategies that consistently exceed regional booking targets. Align with Sales, Marketing, Customer Success, and peer SE managers to share best practices, track market trends, and support organizational scaling. Present on key Identity & Access Management (IAM) and cybersecurity topics at Okta events and industry conferences. Collaborate with HR, Finance, and Operations on workforce planning, budget alignment, and talent acquisition strategies. Serve as an ambassador between the field and headquarters, funneling critical customer feedback and product requirements directly to Okta's Product and Engineering teams. We are looking for demonstrable experience in: A proven track record of managing and scaling high-performing Solution or

awsrestmachine learning
View job →
O
Okta
📍 Washington• Full-time• From $207K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal Operations Engineering Group Okta's Federal Operations team supports government customers operating in FedRAMP-authorized, IL4, and IL5 environments. We deliver the same 99.999% availability promise to the federal market while meeting the strict security, compliance, and operational requirements that come with it. We're looking for a technical leader who understands both the SRE discipline and the unique demands of federal customer relationships — someone who can hold a technical conversation with an agency ISSO in the morning and unblock an incident bridge with an engineering team in the afternoon. As Senior Manager of Federal SRE Operations, you will own the operational health of Okta's federal environments, lead a team of engineers working within compliance-governed change processes, and serve as a trusted point of contact for federal security stakeholders. What you'll be doing Lead

awskubernetesci/cd
View job →
O
Okta
📍 Washington• Full-time• From $189K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Companies all over the world and across every industry are moving to the cloud, and Okta’s identity-centric zero trust security is a key pillar of that digital transformation. Performant and efficient cloud infrastructure management is essential for ensuring the effectiveness of Okta’s services to succeed in every market vertical, geography and segment. With a flexible and resilient platform for cloud infrastructure management, Okta can be a true partner to our customers as they scale and adopt new applications and technology. We are looking for a Sr Product Manager who will be responsible for driving strategy related to Okta’s infrastructure related revenue streams, including examples such as Enhanced Disaster Recovery, Private Cloud, DynamicScale, and our cloud service provider posture. If you are passionate about the intersection of infrastructure and large scale enterprise products, we are looking for you. Responsibilities Develop and execute cloud infrastructure product business cases and strategy that aligns with the company's overall business goals and objectives, as well as broader industry trends and market context Collaborate with cross-functional teams, including engineering, sales, and marketing, to ensure the successful launch and adoption of infrastructure products and services Develop pricing strategies and models for infrastructure products and services Conduct market research and analysis to identify trends and customer needs, and use

awsgcpgit
View job →
🔔

Get new senior infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime