NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
Jobiba hiring network
Senior Infrastructure Automation Engineer Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure automation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What You Will Be Doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuously improve reliability and efficiency. Build and maintain runbooks, standard operating procedures, and comprehensive documentation for storage services and automation.<
Here at Datadog, we think about offensive security a little bit differently. We embrace automation and AI to run adversary simulations continuously across a massive cloud-native environment, and we expect our offensive engineers to build the tooling that makes that possible. We're looking for a Senior Security Engineer who can execute sophisticated red team operations, write the code that scales them, and take an AI-first approach to offensive security engineering. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Plan and execute red team engagements end-to-end, simulating real-world threat actors across cloud infrastructure (AWS, GCP), Kubernetes, CI/CD pipelines, and corporate environments Build and maintain custom offensive tooling, automation frameworks, and engagement infrastructure, treating offensive operations as a software engineering problem Develop custom payloads and evasion capabilities tailored to Datadog's environment and modern defensive controls (EDR, SIEM, network monitoring) Improve the efficiency of offensive operations through thoughtful use of automation and AI, accelerating reconnaissance, vulnerability analysis, and reporting workflows Partner with the Detection & Response team on purple team exercises to validate detection logic, improve alert fidelity, and influence threat models Translate offensive findings into concrete improvements by working directly with defensive security and engineering teams to close gaps Who You Are: You have 5+ years of hands-on experience in offensive security (red teaming, penetration testing, or adversary simulation) with a track record of operating against mature, well-defended environments You write production-quality code (Python, Go, or similar), can build your own tools, and automate your w
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology, anywhere, on any device or app. Our flexible and neutral products, Okta Platform and Auth0 Platform, provide secure access, authentication, and automation, placing identity at the core of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box - we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. The Product Auth0 is a developer-friendly identity platform that simplifies authentication and authorization for applications. Designed by developers for developers, we make access to applications safe, secure, and seamless for the more than 100 million daily logins around the world. Our modern approach to identity enables this Tier-Ø global service to deliver convenience, privacy, and security so customers can focus on innovation. Know more about our product at https://auth0.com/ . The Role We are hiring for a new team within Core Identity, the Engineering organization entrusted with the very heart of the Auth0 application. Our teams own the authentication pipeline, identity protocols, user sessions, and all the fundamental concepts and foundational elements that underpin our entire product. We are also the stewards of the platform's technical he
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Technology, Data and Intelligence Team Message Okta’s Technology, Data and Intelligence (TDI) team delivers the systems, tools, and services that power internal operations across the company. From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will help build, improve, and maintain our cloud platform services by designing and implementing complex cloud-based engineering enablement systems. With a strong focus on automation, testing, and operational excellence, you will deliver foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale. What you'll be doing Secure Cloud Infrastructure & Pipelines: Design, build, and modernize scalable cloud environments and development tools while strictly enforcing security policies and standards for regulated environments. Cross-Functional Collaboration & Advocacy: Partner with software engineering teams to champion DevOps and SRE best practices, deliver excellent internal customer service, and actively contribute to Agile workflows (e.g., demos, architecture sessions). Technical Documentation & Operations: Create and maintain comprehensive technical documentation, including network diagrams, runbooks, and disaster recovery procedures to en
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Team Okta is The World’s Identity Company. We enable everyone to safely use any technology anywhere, on any device or application. Our Workforce and Customer Identity Clouds provide secure, flexible access, authentication, and automation, transforming how people navigate the digital world and placing Identity at the core of business security and growth. At Okta, we are committed to celebrating a variety of perspectives and experiences. We are looking for lifelong learners and individuals whose unique experiences will make our organization better. The Technology Data and Intelligence (TDI) team at Okta is focused on boosting internal efficiency through the implementation of secure, scalable, and innovative systems. About the Role We are seeking an independent Senior Salesforce Engineer who effectively balances technical excellence with a disciplined approach to the software development lifecycle. In this role, you will serve as a guardian of our codebase through implementation, rigorous peer reviews, automated testing, and comprehensive deployment support, ensuring the architectural integrity of our environments. Key Responsibilities Technical Implementation: Lead the translation of user stories into robust, scalable technical designs. Responsible for the end-to-end execution of features, ensuring alignment with organizational standards. Solution Design and Architecture: Architect and design secure, scalable solutions by translating complex
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Team Okta is The World’s Identity Company. We enable everyone to safely use any technology anywhere, on any device or application. Our Workforce and Customer Identity Clouds provide secure, flexible access, authentication, and automation, transforming how people navigate the digital world and placing Identity at the core of business security and growth. At Okta, we are committed to celebrating a variety of perspectives and experiences. We are looking for lifelong learners and individuals whose unique experiences will make our organization better. The Technology Data and Intelligence (TDI) team at Okta is focused on boosting internal efficiency through the implementation of secure, scalable, and innovative systems. About the Role The team is seeking a talented Senior Software Engineer to become part of our TDI team in Bangalore, who can effectively balance technical excellence with a disciplined approach to the software development lifecycle. In this role, you will be responsible for designing and developing customizations, extensions, configurations, and integrations required to meet the company’s strategic business objectives. Candidates will work collaboratively with business stakeholders, business analysts, and engineers on different infrastructure layers, from proposal development to deployment and support. Therefore, a commitment to collaborative problem-solving and delivering high-quality solutions is essential. In a
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Database Reliability Engineer (DBRE) Experience Level: Mid–Senior (4+ years PostgreSQL experience) About the Role We are looking for a highly skilled Database Reliability Engineer (DBRE) with deep expertise in PostgreSQL at scale and solid experience with MySQL. In this role, you will design, operationalize, and optimize the data persistence layer that powers our large-scale, mission-critical systems. You will work closely with SRE, Platform, and Engineering teams to ensure performance, reliability, automation, and operational excellence across our database environment. This is a hands-on engineering role focused on building resilient data infrastructure, not just administering it. Responsibilities: Architecture, Reliability & Performance Design, implement, and operate highly available PostgreSQL clusters (physical replication, logical replication, sharding/partitioning, failover automation). Optimize query performance, indexing strategies, schema design, and storage engines. Perform capacity planning, growth forecasting, and workload modeling. Own high-availability strategies including automatic failover, multi-AZ/multi-region setups, and disaster recovery. Automation & Tooling Develop automation for any and all tasks including but not limited to: provisioning, configuration, backups, failovers, vacuum tuning, and schema management using tools such as Terraform, Ansible, Kubernetes Operators, or custom tooling. Build monitoring, alerti
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Senior Platform Engineer at HP IQ, you will help build and evolve the infrastructure, tooling, and shared platform capabilities that enable our engineering teams to develop and operate reliable, secure, and scalable services across cloud and edge environments . You will work closely with application, services, AI/ML, and security teams to improve developer velocity, production readiness, reliability, and operational efficiency across a heterogeneous infrastructure footprint. What You Might Do Design, build, and maintain shared infrastructure and platform capabilities across cloud and edge environments. Build automation and self-service tooling that improves engineering velocity and operational consistency. Develop and maintain Infrastructure-as-Code, deployment workflows, and environment provisioning. Partner with engineering teams on production readiness, including reliability, security, observability, scalability, and recovery. Improve monitoring, alerting, incident response, and operational tooling across distributed environments. Automate repetitive operational t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Job Overview: We are looking for a Senior Engineer to join the FGA DevEx team and help evolve our end to end developer experience across both OSS and SaaS. This team owns the SDKs in Go, JavaScript, .NET, Python, Java and other languages, along with CLI workflows, IDE integrations, GitHub automation, developer documentation, and release strategy. All development is done in the open as open source, and we actively welcome and review community contributions. Our guiding principle is One developer experience, many deployment models. As a Senior Engineer, you will take ownership of significant portions of the SDK and tooling ecosystem, ensure high quality implementations across languages, and contribute to a consistent and reliable developer experience. Responsibilities: Maintain and enhance existing SDKs for FGA in Go, JavaScript, .NET, Python, and Java, leveraging our SDK generator framework. Customize and refine SDK templates and wrappers to ensure consistency across languages and support configuration overrides such as store ID, authorization model ID, headers, and parallelization limits. Implement and improve core SDK features including client credentials authentication flows, robust error mapping, retry logic with jitter, and rate limiting safeguards. Implement advanced capabilities such as BatchCheck, ListRelations, and non transactional write operations with appropriate parallelization and performance considerations. Contribute to the SDK generator tool
Who are we? FalconX is a pioneering team of operators, investors, and builders committed to revolutionizing institutional access to the crypto markets. Operating at the intersection of traditional finance and cutting-edge technology, FalconX addresses the industry's foremost challenges: Navigating the digital asset market can be complex and fragmented, with limited products and services that support trading strategies, structures, and liquidity found in conventional financial markets. As a comprehensive solution for all digital asset strategies from start to scale, FalconX operates as the connective tissue empowering clients with seamless navigation through the ever- evolving cryptocurrency landscape. Responsibilities Be part of a trading systems engineering team, dedicated to building out the core trading platforms. Work closely with cross functional teams to improve the system reliability, scalability and security. Engage in and improve the quality supporting the platform. Build and manage systems, infrastructure and applications through automation. Provide operational support to internal teams working on the platform. Work on improvements to bring in high efficiency, reduce latency, deploy systems faster. Practice sustainable incident response and blameless postmortems. Together with your engineering team, you will share an on-call rotation and be an escalation contact for service incidents. Implement and maintain rigorous security best practices across all infrastructure, with a focus on minimizing attack surface and ensuring data integrity. Monitor system health and performance with a keen eye for identifying and resolving issues before they affect trading activity. Manage user queries and service requests (often requiring in depth analysis of the technical and/or business logic of our systems). Proactive approach to problem analysis and resolution of production incidents. Manage Issue tracking and prioritisation of day to day production incidents. Manage platf
Join Delphi - Where Innovation meets transformation At Delphi, we believe in creating an environment where our people thrive. Our hybrid work model empowers you to choose where you work—whether it's from the office, your home, or a mix of both—so you can prioritize what matters most. We are committed to supporting your personal goals, family, and overall well-being while driving transformative results for our clients. We welcome exceptional talent from anywhere across the globe. Interviews and onboarding are conducted virtually, reflecting our digital-first mindset. Rooted in the region, we specialize in delivering tailored, impactful solutions in Data, Advanced Analytics and AI, Infrastructure, Cloud Security, and Application Modernization. Whether it’s enabling predictive analytics , transforming operations with automation, or driving customer engagement with intelligent platforms, we are the trusted partner for organizations ready to embrace a smarter, more efficient future. We are looking for a hands-on Senior QA Consultant to lead end-to-end testing of AI and Generative AI applications across enterprise environments. This role combines strong expertise in traditional QA engineering with modern AI evaluation and validation practices. The ideal candidate will drive quality assurance initiatives for RAG pipelines, multi-agent systems, OCR and Speech-to-Text solutions while ensuring production grade quality outcomes in regulated industries such as Finance, Healthcare, and Insurance. The successful candidate will lead a small QA pod, collaborate closely with engineering, AI/ML, product, and business teams, and act as the client-facing QA owner for enterprise AI engagements. This role requires strong technical leadership, automation expertise, AI evaluation capabilities, and excellent stakeholder management skills. Experience Requirements • 7–10 years of experien
Senior Software Engineer (Backend) About Team Cloud Platform Engineering group designs, builds and manages platforms that allow Myntra’s tech product to be secured, reliable, deployed and run at scale. These horizontal platforms leverage cloud hosting platforms and ensure all Myntra’s hostings are agnostic to Cloud vendor. We also build a number of production automation like provisioning of infrastructure at scale and manage complex access management on servers. We have developed numerous in-house tools and platform for load/stress testing, zero trust, security and compliance, CI/CD, Observability at scale. And, we aggressively adopt from open sources and try to contribute back to the community. Tools and Platform Engineering This team builds and maintains centralized and high-scale platforms for Observability (centralized log collection, metric systems, monitoring systems), Security & Compliance (access management, secret management, database access, change management systems, Authentication and Authorization of services etc). The platforms developed by these teams are centralized tools used by all engineering teams for database access, and changes, infrastructure provisioning, on-call scheduling, onboarding new monitoring etc. The vault system and APIs, developed, deployed and managed by this team, is being used by all Myntra’s production services. This team consists of full-stack developers who are skilled in Python, Golang, ReactJs. Roles and Responsibilities Design, build and maintain central platform products to improve the security posture of Myntra Write maintainable, scalable, and efficient code. Design and architect technical solutions for the developer community at Myntra Work in a cross-functional team, collaborating with peers during the entire SDLC. Follow coding standards, code reviews, etc. Follow scrum sprint cycles and commitment to deadlines. Identify security gaps in or for software platforms and incorporate them into requirements
Get new senior infrastructure automation engineer jobs by email
Daily job updates · Unsubscribe anytime