Senior Infrastructure Automation Engineer, Compute Platform — 2 Locations. Apply via Workday.
Jobiba hiring network
Senior Infrastructure Automation Engineer Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure automation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. When was the last occasion you had the opportunity to contribute to a company that is shaping an industry and empowering individuals to translate ideas into tangible impact with speed? Smartsheet's core mission is to empower everyone to enhance their work processes. Our business model is founded on identifying exceptional talent and providing them with the autonomy to develop our acclaimed Software as a Service (SaaS) offering. With a user base exceeding 10 million, our platform is utilized across various industries, including construction, retail, and software development, presenting us with unique technical challenges. Smartsheet is seeking a Senior Business DevOps Engineer to join our Corporate Systems Development team in Bangalore. This role will focus on building and scaling our CI/CD pipelines, infrastructure automation, monitoring frameworks, and deployment processes supporting mission-critical integrations across Finance, People, Sales, Legal, IT, and Engineering systems. You’ll work across a variety of systems and platforms (AWS, GitLab, DataDog, Terraform, Boomi, UiPath) to streamline deployment of backend integrations and automation solutions. If you thrive on optimizing developer velocity, ensuring system reliability, and automating everything from build to deploy, this role is for you. The position reports to the Senior Manager, Systems Development and collaborates closely with global developers, architects, and application administrators to ensure our platform foundations are secure, efficient, and scalable
Senior Infrastructure Architect — Enterprise Observability and Automation Description - Job Summary Senior individual contributor responsible for the architecture, implementation, and operational ownership of enterprise observability, monitoring, and automation platforms across HP's global IT environment. This role modernizes infrastructure visibility capabilities while ensuring operational stability, security, and compliance. Serves as a technical and operational bridge between infrastructure engineering, cybersecurity, SOX/compliance stakeholders, automation teams, and external technology partners — leading complex initiatives such as platform migrations, enterprise integrations, and governance enablement. Responsibilities Enterprise Observability and Monitoring Application owner and senior technical authority for enterprise monitoring and logging platforms (Datadog, Splunk), including platform governance, roadmap alignment, and operational oversight. Lead enterprise-scale monitoring platform migrations, including architecture design, agent strategy, data ingestion models, vendor coordination, and deployment across 5,000+ servers. Define standards for alerting, dashboards, observability data quality, and integration with ITSM platforms (ServiceNow). Design and manage multi-org Datadog architecture, including org structure, RBAC, SSO/SAML, secrets management, and cybersecurity compliance. Oversee SNMP-based monitoring of storage and network devices, including device profiling, syslog/event integration, and NetFlow collection. SOX Compliance and IT Governance SOX control owner for enterprise monitoring applications — approve monthly reviews, participate in internal/external audits (EY), and maintain ITGC/SOX compliance. Provide audit evidence, walkthrough docu
Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As the Senior Software Engineer, Tooling and Development Infrastructure, you will play a critical role in shaping the developer productivity tools and automated testing strategy. You’ll collaborate closely with design, development, and quality teams to plan, design, and implement robust automated tools and services that ensure the quality and reliability of our AI software stack. You will be highly hands-on in your work and collaborate closely with stakeholders. This position offers a unique opportunity to influence the development of cutting-edge automation frameworks, foster a culture of quality, and contribute to the long-term success of the organization. What You Might Do Develop and implement automation frameworks and testing strategies that cover the entire software stack, from backend systems to user-facing features. Identify, evaluate, and integrate new tools that streamline development. This includes everything from code quality tools and to Infrastructure-as-Code (IaC) solutions. Lead continuous improvement efforts for our build, release, and test systems, ensuring a robust
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we are building the future of secure, enterprise-grade cloud automation and system connectivity. We are looking for a Senior Software Engineer to join our global Automation Engineering team to design, scale, and govern our enterprise integration substrate using AWS cloud services and modern iPaaS platforms. This is a senior individual contributor role for a hands-on system engineer who can set technical standards, establish integration best practices, and partner with cross-functional teams to automate complex business workflows at scale. What You'll do : Lead technical design and execution for automation initiatives within the team, creating paved paths that enable builders across Okta to connect enterprise systems seamlessly. Design, build, and deploy high-throughput event-driven integration flows , API gateways, and async orchestration workflows using AWS architectures and iPaaS platforms. Build reusable frameworks , developer SDKs, self-service primitives, and integration templates to streamline automation delivery. Serve as a technical domain expert on AWS cloud services and modern iPaaS tooling, driving scalable architecture, reliability, and builder enablement. Partner with operations, security, and platform teams to strengthen monitoring, observability, structured audit logging, and automated governance. Mentor team members and participate in cross-functional design reviews , instilling a platform engineering mindset and raising the b
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance. You will: Design and Implement Observability Solutions : Develop comprehensive monitoring and alerting systems using modern observability tools. Create dashboards and metrics that provide real-time visibility into system health and performance. Implement logging strategies that enable quick problem identification and resolution. Drive Automation and Infrastructure as Code : Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi. Design and maintain CI/CD pipelines that enable reliable and consistent deployments. Create self-healing systems that can automatically respond to common failure scenarios. Establish SLOs and SLIs : Work with product and engineering teams to define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to track and report on these metrics, ensuring we maintain high reliability standards while balancing innovation speed. Incident Management and Response : Lead incident response efforts, conducting thorough post-morte
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are looking for a Senior Automation Engineer to design and scale intelligent automation across Okta’s Enterprise Technology ecosystem. This role will reduce manual operational work, improve developer productivity, and accelerate the software delivery lifecycle across platforms such as Salesforce, NetSuite, AEM, Boomi, identity systems, CI/CD platforms, and other enterprise applications. The engineer will build reliable automation services, reusable integration patterns, and AI-assisted developer agents that understand source code, metadata, configuration, tickets, documentation, deployment history, and operational context. These solutions will automate release activities, identify deployment risks, reduce merge-conflict resolution effort, improve environment management, and enable safe self-service operations. The role requires a combination of software engineering, enterprise application expertise, API and integration development, CI/CD automation, security engineering, and practical experience applying AI to developer and operational workflows. The ideal candidate is comfortable moving between platform-specific technologies and building scalable, secure, platform-agnostic automation frameworks. What you'll be doing Intelligent Automation and Developer Productivity Build AI-assisted developer agents and automation services that reduce repetitive engineering and operational work. Use contextual information from Git repos
Work Flexibility: Remote As a Senior Lead, Data Engineering, you will serve as a technical leader who helps shape the future of enterprise data solutions. In this role, you will drive complex data initiatives, influence technical strategy, and partner with teams across the organization to build scalable, high-impact data products. This is an opportunity to solve challenging business problems while mentoring fellow engineers and elevating data engineering best practices. What You Will Do Lead the architecture, development, and modernization of scalable enterprise data platforms that support global procurement analytics and business transformation. Define and help execute a multi-year data engineering strategy focused on platform scalability, reliability, automation, technical debt reduction, and long-term maintainability. Design, build, and optimize Azure-based data solutions using technologies such as Databricks, Delta Lake, Azure Data Factory, Azure DevOps, CI/CD pipelines, and infrastructure automation. Integrate and harmonize data across multiple ERP systems by standardizing supplier, purchasing, and master data into common enterprise data models. Partner with procurement analysts, architects, engineers, and business stakeholders to translate complex business needs into reusable, scalable data products and engineering solutions. Establish engineering standards, conduct architecture reviews, improve documentation, and mentor engineers to raise the overall technical capability of the team. Identify and implement AI-enabled approaches that accelerate development, improve data quality, automate documentation, support testing, and enhance analyst productivity. Evaluate and recommend tools, frameworks, patterns, and platform investments that improve performance, reliability, security, governance, and operational ef
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Roles & Responsibilities Design, deploy, and manage AWS cloud infrastructure and Kubernetes (EKS) environments. Build, maintain, and optimize GitLab CI/CD pipelines for automated deployments. Implement DevSecOps best practices, including infrastructure security, vulnerability management, and compliance controls. Monitor, troubleshoot, and improve platform reliability, performance, and scalability. Deploy and manage cloud-native applications, databases, and supporting services. Collaborate with development teams on architecture, deployment strategies, and production readiness. Drive infrastructure automation using Infrastructure as Code (Terraform). Support incident management, root cause analysis, and operational excellence initiatives. Mandatory Requirements 5+ years of hands-on DevOps/DevSecOps experience in production environments. Strong expertise with AWS services including EKS, EC2, S3, CloudFront, WAF, Lambda, and VPC. Minimum 3+ years of Kubernetes experience, preferably Amazon EKS. Hands-on experience building and managing GitLab CI/CD pipelines. Strong Linux administration and troubleshooting
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we are building the future of secure, enterprise-grade cloud automation and system connectivity. We are looking for a Software Engineer to join our global Automation Engineering team to design, scale, and govern our enterprise integration substrate using AWS cloud services and modern iPaaS platforms. This is a individual contributor role for a hands-on system engineer who executes moderately complex tasks, builds platform components and collaborates under senior guidance. What You'll do : Contribute technical design and execution for automation initiatives within the team, creating paved paths that enable builders across Okta to connect enterprise systems seamlessly. Design, build, and deploy high-throughput event-driven integration flows , API gateways, and async orchestration workflows using AWS architectures and iPaaS platforms. Build reusable frameworks , developer SDKs, self-service primitives, and integration templates to streamline automation delivery. Develop and maintain AWS cloud services and modern iPaaS tooling, driving scalable architecture, reliability, and builder enablement. Partner with operations, security, and platform teams to strengthen monitoring, observability, structured audit logging, and automated governance. Actively participate in code reviews, team agile ceremonies, and technical discussions Collaborate with security, compliance, and business stakeholders to ensure self-service automations are secure, resilient, zero-tr
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Software Engineer, Platform Engineering (L2) About the job This position is an engineering role within Twilio Platform Engineering, suited to an engineer who is building hands-on experience with highly available, large-scale distributed systems. Our systems regularly process more than 12 billion emails during peak events like Black Friday, and our throughput requirements continue to scale rapidly. As an L2 engineer, you will help build and operate backend services at scale, working alongside more senior engineers on our dual-cloud infrastructure spanning Amazon Web Services (AWS) and Microsoft Azure. You'll get hands-on with Kubernetes, contribute to Terraform-based infrastructure automation, and write production code to help keep distributed systems healthy under real production load, using modern AI-assisted tooling to move faster and ramp up your skills. Responsibilities In this role, you’ll: WEAR THE CU
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure. This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience
Get new senior infrastructure automation engineer jobs by email
Daily job updates · Unsubscribe anytime