Jobiba hiring network

Senior Infrastructure Automation Engineer Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior infrastructure automation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

NVIDIA is well positioned as the 'AI Computing Company', our GPUs being the brains that power modern Deep Learning software frameworks, accelerated analytics, modern data centers, and driving autonomous vehicles. We are looking for a Senior Software QA Test Development Engineer to join in the mission of crafting a distributed technology for all NVIDIA teams that remotely manage 10s of 1000s of resources in a simple and controlled fashion, allowing engineers to focus on engineering and automation, rather than being burdened by manual operational tasks. SWQA test developer engineers at NVIDIA are responsible for creating test plans, execution, and reporting, as well as developing scripts for test automation, designing and developing tools for the QA team, and developing integration tests for validation. As a test developer, you must identify weak spots and constantly design better and more creative test plans to break software and identify potential issues. You will have a huge impact on the quality of NVIDIA's products. The ideal candidate must have strong programming skills and hands-on experience using AI development tools to improve quality and productivity across the end-to-end QA workflow. This includes leveraging AI assistants for test automation, code generation, debugging, and enhancing testing efficiency. During the interview process, we will assess your ability to effectively use AI development tools and evaluate your programming capabilities to ensure you can deliver high-quality solutions. What you’ll be doing: Architect, implement, and evolve scalable agentic end to end SWQA workflow, automated test frameworks, infrastructure, and tooling for complex software products. Define test strategy and quality gates across functional, integration, regression, reliability, and release-validation workflows. Build and maintain high-value automated coverage for Linux-based, co

dockerkuberneteslinux
View job →
R
17 days ago

Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. 5 Work Days Per Week Office at Tai Seng Exchange Tower B Near Tai Seng MRT, Singapore Insurance Coverage Entitled to Yearly Bonus & Performance Bonus Responsibilities (Site Reliability Engineer - Senior / Lead ) High Availability and Stability Maintenance of Application Systems: Includes daily monitoring, alert response, emergency handling, on-call duties, regular system health checks, and performance optimization. Compliance and Secure Access Construction for Application Systems: Ensure operational design, processes, and data management comply with relevant privacy and data protection laws. Ensure compliance with full auditing and regulatory checks and provide auditing materials as required. Change and Release Management: Best practices for application system changes, including change control, version management, and rollback strategies, while ensuring operational duties during release windows. Automation and Infrastructure Optimization: Drive operational automation by designing and implementing automated tools and processes, ensuring resource allocation is optimized and supporting business scalability. Other Operational Practices and Work Arrangements: Providefeedback and suggestions for business architecture design and continuously produce operational technical documentation. Qualifications Bachelor's Degree or above; a degree in computer science or a related field is preferred. Experiences as Senior SRE or

awsazuredocker
View job →

About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr

pythonawsgcp
View job →
R
Roblox
📍 San Mateo• Full-time• From $196.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. At Roblox, we strive to connect a billion people with optimism and civility, and the Safety organization’s mission is to become the leader in civil immersive online communities. We systematically detect, remove, and prevent problematic accounts, content and behavior, and we make Roblox accounts secure and free from compromise. We cover a broad area of the tech spectrum, including machine learning, classifiers for 3D models, experimentation, automation, detection workflows, and AI-powered text filters. Aligned and partnering with product teams, we use this tool-belt to discover new opportunities, influence and shape the product roadmap and prioritization, build safety products, and measure the impact on our community of users and developers. In doing so, we keep Roblox safe, civil, and inclusive, and we foster positive relationships between people around the world. WHY Safety Data Infrastructure To pro-actively find bad actors and protect good users, Roblox needs to ingest enormous amounts of data and make it usable both to our automated detection systems and for human moderators and customer support agents doing investigations. The Safety Data Infrastructure team is addressing these needs w

pythonawsgit
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the Platform Core Automation team, you'll build the AI infrastructure and agentic systems that automate customer support and compliance operations at Coinbase. This team is reimagining what these processes look like when AI handles them end-to-end, replacing manual workflows with intelligent agents that resolve customer issues and execute compliance tasks autonomously. You'll own the design and delivery of LLM-powered systems, grounding techniques, and integration pipelines that directly reduce resolution times, cut costs, and improve accuracy across millions of customer interactions. What you’ll be doing (ie. job duties): Own the design and delivery of agentic AI systems that power Coinbase's customer support and compliance automation, from LLM orchestration through production deployment and monitoring. Build scalable, secure backend infrastructure in Python and Golang that serves AI workloads, including model integration pipelines, guardrails, grounding mechanisms, and measurement frameworks. Drive end-to-end project execution on complex AI initiatives, making technical trade-offs across latency, accuracy, cost, and reliability. Partner with Customer Experience, Compliance, and product engineering teams to identify high-impact automation opportunities and translate operational pain points into AI-powered solutions. Strengthen engine

REMOTEpythonmongodbaws
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Software Engineer, AI Transformation You'll join a high-performing team of engineers driving AI transformation at Coinbase as a Senior Software Engineer on the IT Operations team within ESTO. This team builds custom products through full-stack engineering and scales the infrastructure powering Coinbase's AI products, with direct exposure to senior leadership in a fast-paced, incubator-style environment. You'll own the reliability and automation of critical AI infrastructure, ensuring our systems are resilient, observable, and secure at scale. What you'll do: Own end-to-end delivery of AI products by building production-grade distributed systems, including serving infrastructure, data pipelines, and deployment orchestration across the full stack throughout the SDLC. Drive platform adoption by designing clean APIs, abstractions, and developer-facing tooling that enable product teams to integrate AI capabilities without bespoke infrastructure requests. Partner with engineering and product leadership across Platform and other product groups to align infrastructure requirements, resolve cross-team technical dependencies, and define shared platform contracts. Shape engineering standards and technical culture by establishing architectural patterns, mentoring engineers, and raising the bar on code quality, observability, and operational excellence. Build f

REMOTEpythonawsdocker
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $253.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Core Automation team envisions a future where AI-powered operations lead to seamless, delightful experiences for customers and maximum value for shareholders. Towards this vision, our strategy is to 1) start with re-imagining and automating Compliance operations with Agentic AI, 2) incorporate learnings from this automation experience and build primitives and orchestration solutions that allow us to scale, and 3) scale across a variety of domains enabling AI to enhance every part of customer experience and internal processes. It is day 1 for us. We are laser focused on automating compliance processes today. This is not just focused on driving efficiencies by automating manual processes, but reimagining what processes and systems must look like in a fully AI driven automated world, and then bring that vision to reality, working across cross functional teams to offer a delightful customer experience, improve compliance posture and deliver outsized shareholder value. What you’ll be doing (ie. job duties): Explore and apply advanced GenAI techniques, including large language models (LLMs) and Agentic AI, to solve complex challenges across the organization. Partner with and influence Operations and Compliance organizations to maximize impact. Build and manage engineering teams in the Compliance domain, to guide the development of features, services, and infrastructure

REMOTEawsaigo
View job →
M
Mongodb
📍 Austin; New York City; San Francisco; Seattle; United States• Full-time• From $127K/yr
1mo ago

We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs. The InfraSec team collaborates closely with other engineering teams to ensure that our infrastructure adheres to the highest security standards. They build essential security infrastructure and implement controls that reinforce the platform’s security posture. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions.This team is deeply involved in the technical aspects of security and the nuances of its actual implementation. This role can sit in our New York City, Austin, Seattle or San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. Responsibilities: Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud platforms (AWS, Azure, GCP), including network and compute security, identity management, and cloud security posture management (CSPM) Automation and Monitoring: Build automated solutions for real-time security monitoring, logging, and alerting in cloud environments. Leverage native cloud services and third-party tools for runtime security monitoring and anomaly detection Security Tooling: Evaluate, implement, and manage cloud-native security tools and platforms for endpoint security, identity management (IAM), and CSPM Qualifications: Experience: 6+ years of experience in SRE, infrastructure engineering or similar role, with a strong focus on security work, with ideally 2+ years in a senior or staff engineering role Security Mindset: A comprehensive understanding of all facets of cloud environment security, spanning from foundational OS networking laye

mongodbawsazure
View job →
O
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Infrastructure Platform and Shared Services Team Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling. As the Sr. Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, Observability, automation platform & tooling. What you’ll be doing Lead the Infra platform and shared services org and various initiatives across SRE & Infrastructure organization. Build a world-class observability platform and monitoring capabilities enabled with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and intuitive self-service capabilities. Own the design and operation of scalable, self-service Cloud infrastructure platforms (e.g. Observability Platform, SRE Productivity, deployments, and Edge Infrastructure) Lead, mentor, and grow a high-performing team of engineers and managers across SRE and infrastructure shared services domains. Perform engineering design evaluations and ensure the completion of projects within resource,

awsci/cdrest
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxartificial intelligence
View job →
G
17 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxai
View job →
A
Affirm
📍 Poland• Full-time• Remote• $384K – $576K/yr
17 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Spain• Full-time• Remote• From €1M/yr
17 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →
G
GHX
📍 Hyderabad• Full-time
17 days ago

Role: Senior Cloud Engineer Location: Hyderabad, India (Hybrid) Department: Product Development About GHX: GHX (Global Healthcare Exchange) is a leading healthcare technology company on a mission to simplify the business of healthcare and improve patient outcomes. Founded in 2000, GHX has built the GHX Global Network — the world’s largest cloud-based supply chain community connecting healthcare providers, suppliers, distributors, and partners to automate key processes, reduce costs, and increase operational efficiency. Its solutions span electronic trading, procurement automation, inventory and contract management, business intelligence, and data synchronization, helping healthcare organizations improve productivity and focus more on patient care. Over the years, GHX has enabled significant cost savings for the industry and continues to innovate with intelligent automation and AI-driven capabilities. Website: https://www.ghx.com/ LinkedIn: https://www.linkedin.com/company/ghx/ About Role: The Senior Cloud Engineer leads the design, implementation, and operations of the organization’s cloud infrastructure. This role is responsible for complex projects, high-level architectural planning, and making strategic technology decisions that support scalability, performance, security, and cost optimization. Acting as a technical leader, the Senior Cloud Engineer provides mentorship to junior and mid-level engineers while driving innovation, resiliency, and compliance in cloud environments. Key Responsibilities: Design and Implementation Lead the design and implementation of cloud-native architectures and hybrid cloud solutions. Oversee cloud engineering projects, ensuring solutions are secure, resilient, and cost-effective. Contribute to high-level architectural planning and strategic decision-making around technology adoption. Review and approve design proposals, Infrastructure as Code (IaC) tem

awsazuregcp
View job →
O
Okta
📍 San Francisco• Full-time• From $165K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en

pythonsqlpostgresql
View job →
🔔

Get new senior infrastructure automation engineer jobs by email

Daily job updates · Unsubscribe anytime