Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp
Jobiba hiring network
Senior Infrastructure Automation Engineer Jobs
7,101 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure automation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At NVIDIA, we are at the forefront of technological innovation, pushing the boundaries of AI and accelerated computing. Our team in Santa Clara, CA is looking for a Senior Software Engineer in Test to join us in this exciting journey. This is a ground breaking opportunity to work with powerful technology, collaborate with a world-class team, and make a significant impact in the industry. If you are passionate about AI and quality assurance, and thrive in a dynamic environment, this role is perfect for you! What you'll be doing: Accomplishing test cases to validate NVIDIA enterprise offerings, such as NIM, NeMo, and BioNeMo. Crafting, implementing, and maintaining automated test cases and supporting automation infrastructure. Collaborating with development teams to triage issues, perform root cause analysis, verify fixes, define additional tests, and improve test plans. Investigating and bringing to bear AI capabilities to accelerate the Quality Assurance (QA) process. What we need to see: MS or PhD degree in computer science or relevant field, or equivalent experience. At least 5+ years of professional experience in software testing. Proficiency in oral and written English. Comfort working with Linux OS. Strong skills in shell and Python programming. Strong knowledge of QA principles and background in software testing. Experience using AI development tools for crafting test plans, developing test cases, and automating test cases. Excellent problem-solving abilities. Strong interpersonal skills, quick learning ability, proactive approach, innovation, and dedication. Self-motivation and a passion for learning new hardcore technology. Knowledge in LLM and AI models is a plus. <
The Engineering Lead Analyst – Test Automation Platform Engineering is a senior-level technical leadership role responsible for driving the architecture, implementation, management, operational support, and continuous enhancement of enterprise test management and automation platforms. In this role, you will lead efforts to modernize test automation capabilities across the global technology ecosystem. You will architect end-to-end integration workflows, embed automated quality gates into enterprise CI/CD pipelines, and administer as well as operationally support both vendor and internally developed enterprise platforms (e.g., Core Performance Engineering / Performance Center, ALM-Quality Center, Zephyr Enterprise, CSDP / Octane). Additionally, you will play a critical role in production operations—delivering tier-3 platform support to rapidly and safely troubleshoot, triage, and remediate performance and availability issues in complex, distributed production environments. The ideal candidate blends deep hands-on expertise in software testing frameworks, modern DevOps pipelines, containerized infrastructure, message-driven integration, and robust operational resilience practices with strong governance, compliance, and stakeholder leadership skills. Key Responsibilities 1. Platform Engineering & Operational Support Install, configure, upgrade, administer, and support enterprise test management and performance engineering toolsets (e.g., Core Performance Engineering / Performance Center, ALM-Quality Center, Zephyr Enterprise, Core Software Development Platform [CSDP] / Octane). Provide end-to-end operational support for
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role We're seeking a Senior/Staff Engineer to build and maintain the automation infrastructure that powers the development cycles of our North platform. This engineer will design and implement robust automation systems that enable engineers to efficiently test and validate changes across diverse environments and configurations. This role sits at the intersection of infrastructure and standards. You'll build the systems, frameworks, and culture that allow the rest of engineering to own quality themselves; improving and extending our testing platform by creating the infrastructure that allows engineers to write and execute tests, and enable every engineering team to ship with more confidence. Key Responsibilities Design and implement automation pipelines that support comprehensive testing across multiple environments with varying feature flags and realistic customer data profiles Create intelligent testing agents that simulate real user behavior to validate different configuration combinations Develop and maintain GitHub workflows and actions to automate testing, deployment, and validation processes Manage and optimize H
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will help develop and validate next-generation server platforms that power a reliable, high-performing, and cost-efficient infrastructure at scale. You will work across platform bring-up, firmware qualification, hardware validation, fleet integration, and performance optimization to support large-scale production deployments. You Will: Bring-up & Sustaining: Drive key aspects of the hardware development lifecycle, including feasibility studies, hardware bring-up, validation, deployment, and ongoing production support. Platform Optimization: Perform platform integration, performance characterization, and system-level debugging across compute infrastructure, focusing on hardware optimization, driver tuning, and thermal/power efficiency. Hardware Validation: Develop and execute rigorous evaluation and stress-testing strategies for server platforms to ensure reliability and performance under production-scale workloads. Firmware & Fleet Enablement: Support BIOS/BMC firmware qualification, hardware health monitoring, and automation tooling for firmware deployment and lifecycle management. Vendor & Cross-Functi
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our New York City, Austin, Seattle or San Francisco offices, or work fully remotely on standard East Coast business hours. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls and services within MongoDB Atlas multi-cloud infrastructure. The team sits within the Site Reliability Engineering organization and works with other engineering teams to ensure that our infrastructure adheres to the highest security standards. This role can be based out of our Dublin office, or work fully remotely in Ireland. Responsibilities: Design and build core security primitives and services that protect MongoDB Atlas compute, networking, and identity across AWS, Azure, and GCP Build secure-by-default infrastructure using Linux security mechanisms (AppArmor, SELinux, seccomp, cgroups), Kubernetes, and eBPF to enforce runtime policies and gain deep visibility into systems behaviour Develop APIs, automation, and tooling that manage security posture at scale (CSPM, vulnerability management, workload identity) and provide monitoring, logging, and alerting pipelines that integrate with our tooling (Grafana, Splunk, Victoria Metrics.) Integrate security into our CI/CD and infrastructure-as-code workflows (Terraform) so that security controls are versioned, reviewed, and deployed just like any other code Lead complex projects end‑to‑end, from problem discovery and design docs to implementation, rollout, and long‑term ownership Collaborate with SRE, platform and product engineering teams to define secure architectures for new infrastructure and services Qualifications: You might be a great fit if you match some of the following: 5+ years of experience in Software Engineering, Site Reliability Engineering, or similar roles, preferably with relevant security work Proficiency with at least one programming language (Java, Golang, Rust, Python, or C/C++) and experience with infrastructure-as-code tools (Terraform) to automate security configurations and processes A deep understanding of Linux and networking concepts,
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Security Operations Engineer at Brex, you will focus on preventing, detecting and responding to security threats across Brex's corporate and cloud environments. You will use existing systems and develop tools to improve our security capabilities. Our team is responsible for functions across corporate security, detection & response and infrastructure security domains; and we perform systems engineering and automation to support those functions. Security Operations is part of our wider Trust & IT o
Job Summary Reporting to the Memory Validation leadership team, the Senior Silicon DDR/HBM Validation Engineer will be responsible for the bring-up, validation, characterization and debug of advanced memory subsystems used in next-generation AI compute platforms. The role will focus on DDR and HBM technologies, working closely with silicon design, firmware, characterization, platform and systems teams to ensure robust memory subsystem functionality, performance and reliability. The successful candidate will take ownership of significant validation activities, contribute to debug and root-cause analysis efforts, and help improve validation methodologies, automation and infrastructure. The Team The Memory Validation team sits within the Validation organisation and is responsible for the bring-up, validation, characterization and debug of memory subsystems across Graphcore silicon and platform products. The team supports DDR and HBM validation activities throughout the product lifecycle, from first silicon through production readiness. Engineers work closely with architecture, RTL, firmware, characterization, systems and platform teams to ensure memory technologies meet functionality, performance, reliability and performance objectives. Responsibilities and Duties Execute validation and bring-up activities for DDR and HBM memory subsystems Verify memory bring-up software, firmware and scripts against defined project requirements Debug firmware, hardware and system-level issues and contribute to root-cause analysis activities Analyse system logs, validation data and characterization results to identify failures and performance issues Perform PHY characterization and analog-level analysis during stress testing and validation activities Develop and execute functional, stress, performance and corner-case validation tests Perform signal integrity, voltage, frequency and timing measurements using laboratory instrumentation Char
Location Details: Remote, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... We are looking for an experienced engineer to join the Infrastructure Management and Automation team. In this role, you’ll help shape and build the tools, platforms, and automation that power and manage GoDaddy’s growing infrastructure footprint while leading the design of scalable solutions to complex operational challenges. Our infrastructure platform empowers GoDaddy engineers to help small business customers start, grow, and run their ventures. GoDaddy teams rely on our infrastructure solutions every day, whether for development workflows or hosting production workloads at global scale. If you enjoy solving complex technical problems, designing scalable systems, influencing technical direction, and improving how infrastructure is managed and automated, we’d love to hear from you! What you'll get to do... Lead the design and evolution of GoDaddy’s next generation of infrastructure management platforms and automation systems, taking ownership of complex technical areas from design through delivery and operation Leverage modern engineering practices and AI-assisted development workflows to rapidly design, build, and deliver scalable internal tooling and automation Design and build reliable services and integrations across infrastructure platforms and operational systems, including provisioning, networking, IPAM, DNS, CMDB, and asset management workflows Partner with engineering and operations teams to identify opportunities, shape technical solutions, and improve infrastructure reliability, scalability, utilization, and opera
About Prophecy The leader in AI-native data preparation and analysis, Prophecy is revolutionizing how the world’s top enterprises turn data chaos into reliable insights. We introduce the AI-native data lifecycle (generate, refine, deploy) where our industry leading AI agents and humans work hand-in-hand in visual and document interfaces to analyze, transform and prepare data, to ship trusted insights at enterprise scale. Don’t miss the rocket ship—join Prophecy and build the next data revolution. Position Summary This is a high-impact opportunity to be a senior DevOps engineer in a fast-growing startup, based in Prophecy’s India engineering center. You will own and evolve the foundations that keep our engineering org fast, secure, and cost-efficient at scale — spanning cloud cost management, security DevOps, and CI/CD and engineering operations. You will work with a team of dynamic engineers who take pride in solving complex problems, and you will have the autonomy to set direction in your areas of ownership. The Impact You Will Have Cloud cost (FinOps) Own and evolve our cloud cost optimization program across AWS, Azure, GCP, Databricks, Snowflake and Bigquery building on the programmatic monitoring and controls Analyze billing, asset-inventory, and utilization data across departments to identify wasteful spend and provide actionable insights to optimize it. Develop and maintain cost optimization strategies, roadmaps, and forecasting models Partner with engineering teams to design cost-efficient architectures without compromising scalability or reliability. Build and maintain automation for infrastructure provisioning, scaling, and cost control. Security DevOps Partner with engineering and security to drive our security-hardening program across workstreams such as identity & access governance, secrets & credential lifecycle, cloud access, and CI/CD hardening. Implement and automate guardrails: secrets management, least-privilege access, cr
Who We Are At Justworks, you’ll enjoy a welcoming and casual environment, great benefits, wellness program offerings, company retreats, and the ability to interact with and learn from leaders in the startup community. We work hard and care about our most prized asset - our people. We’re helping businesses get off the ground by enabling them to focus on running their business. We solve HR issues. We’re data-driven and never stop iterating. If you’d like to work in a supportive, entrepreneurial environment, are interested in building something meaningful and having fun while doing it, we’d love to hear from you. We're united by shared goals and shared motivations at Justworks. These are best summed up in our company values, which are reflected in our product and in our team. Our Values If this sounds like you, you’ll fit right in. Department Platform Engineering Who You Are You are a tooling-focused engineer who obsesses over developer productivity. You’ve built and maintained CI/CD pipelines, developer tooling, and automation that makes engineering teams faster. You understand that the best developer tools disappear into the background - they just work. You’re equally comfortable debugging a flaky GitHub Actions workflow and designing a new internal CLI. You care deeply about reducing friction and cognitive load for your fellow engineers. This is a foundational role. You’ll be one of the first dedicated engineers on a platform organization supporting 300+ engineers. You’ll shape not just the technical foundations but the culture and practices of the team. Your Success Profile What You Will Work On Own and evolve CI/CD infrastructure (GitHub Actions, deployment pipelines) to improve build times and deployment frequency Build and maintain our developer portal (Backstage) - service catalog, golden path templates, and TechDocs integration Create golden path templates for Go, Rails, and Vue.js services that bake in best practices automatically Build and maintain internal
Get new senior infrastructure automation engineer jobs by email
Daily job updates · Unsubscribe anytime