Jobiba hiring network

Senior Infrastructure Architect Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior infrastructure architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.

SC
Sigma Computing
📍 San Francisco• Full-time• $150K – $240K/yr
16 days ago

About the Role Sigma is transforming how businesses run by delivering a high performance platform on the modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing Solve challenging problems that arise in providing high performance interactive experience to enable analytics and workflows use cases on top of modern warehouses Build software using the latest developer tools and using programming languages like Rust, Go, GraphQL, Typescript Develop new algorithms and techniques for improving the performance and interactivity for enabling analytics and workflows for the world largest companies Triage product or system issues and debug/track/resolve by analyzing the sources of issues Design and implement new software features to support our fast growing user growth Collaborate with cross-functional groups - infrastructure, design, product, customer support, sales and marketing to build an innovative product capabilities Qualifications We Need 5+ years industry experience building and maintaining high-quality software Experience building and deploying robust and secure web applications in a continuous deployment environment Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong Computer Science fundamentals Qualifications We Want (also, skills you’ll learn!) Experience building software capabilities for analyzing large scale data web applications Data driven aptitude and its application to solve distributed system problems Data model design, and API development experience to enable customer facin

typescriptpythonsql
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $170K – $240K/yr
16 days ago

About the Role Sigma is transforming how businesses run by delivering a high performance platform on the modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing Solve challenging problems that arise in providing an interactive experience on data warehouses for data exploration and analysis Build with modern tools and languages like Rust, Go, GraphQL, Node, and Kubernetes Build backend distributed services, new algorithms and modern API to support a cloud application Triage product or system issues and debug/track/resolve by analyzing the sources of issues Design and implement new software features to support our fast growing user base Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies Qualifications We Need 5+ years industry experience building and maintaining high-quality software Experience building and deploying robust and secure web applications in a continuous deployment environment Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong Computer Science fundamentals Qualifications We Want (also, skills you’ll learn!) Data driven aptitude and its application to solve distributed system problems Data model design, and API development experience SQL query optimization and database internals Administered cloud service infrastructure (GCP, AWS, Azure) Prior experience working at high growth company solving technical problems to enable continued success Additional Job details The base salary range for this position is $170k - $240

pythonsqlaws
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $288K/yr
16 days ago

About Scale AI Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. Role Overview As a Senior Staff Frontier Agents Engineer on our Enterprise team, you'll be the technical bridge between Scale AI's cutting-edge AI capabilities and our most strategic customers. You'll work with enterprise clients to understand their unique challenges, architect custom AI solutions, and ensure successful deployment and adoption of AI systems in production environments. This is a hands-on technical role that combines deep engineering expertise with customer-facing problem solving. You'll work directly with customer engineering teams to integrate AI into their critical workflows. Key Responsibilities Customer Integration & Deployment Partner directly with enterprise customers to understand their technical infrastructure, data pipelines, and business requirements Design and implement custom integrations between Scale AI's platform and customer data environments (cloud platforms, data warehouses, internal APIs) Build robust data connectors and ETL pipelines to ingest, process, and prepare customer data for AI workflows Deploy and configure AI models and agents within customer security and compliance boundaries AI Agent Development Develop production-grade AI agents tailored to customer use cases across domains like customer support, data analysis, content generation, and workflow automation Architect multi-agent systems that orchestrate between different models, tools, and data sources Implement evaluation frameworks to measure agent performance and iterate toward business objectives Design human-in-the-loop workflows and feedback mechanisms for continuous agent improvement Prompt Engineering & Optimization Create sophisticate

pythonawsazure
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $216K/yr
16 days ago

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai

awsazuregcp
View job →
G
16 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxai
View job →
DU
16 days ago

About the Team Data is at the foundation of DoorDash success. The Data Engineering team builds database solutions for various use cases including reporting, product analytics, marketing optimization and financial reporting. By implementing pipelines, data structures, and data warehouse architectures; this team serves as the foundation for decision-making at DoorDash. About the Role DoorDash is looking for a Senior Data Engineer to be a technical powerhouse to help us scale our data infrastructure, automation and tools to meet growing business needs. This is a hybrid position and you must be located in Sunnyvale, San Francisco, or Seattle. You're excited about this opportunity because you will... Work with business partners and stakeholders to understand data requirements Work with engineering, product teams and 3rd parties to collect required data Design, develop and implement large scale, high volume, high performance data models and pipelines for Data Lake and Data Warehouse Develop and implement data quality checks, conduct QA and implement monitoring routines Improve the reliability and scalability of our ETL processes Manage a portfolio of data products that deliver high-quality, trustworthy data Help onboard and support other engineers as they join the team We're excited about you because... 5+ years of professional experience 3+ years experience working in data engineering, business intelligence, or a similar role Proficiency in programming languages such as Python/Java 3+ years of experience in ETL orchestration and workflow management tools like Airflow, Flink, Oozie and Azkaban using AWS/GCP Expert in Database fundamentals, SQL and distributed computing 3+ years of experience with the Distributed data/similar ecosystem (Spark, Hive, Druid, Presto) and streaming technologies such as Kafka/Flink. Experience working with Snowflake, Redshift, PostgreSQL and/or other DBMS platforms Excellent communication skills and experience working with technical and non-tec

pythonjavasql
View job →
G
Guidepoint
📍 Toronto• Full-time• C$135K – C$210K/yr
16 days ago

Overview: Guidepoint seeks an experienced Data/AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The Senior AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization

pythonreactaws
View job →
O
Okta
📍 Bengaluru• Full-time
19 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we are building the future of secure, enterprise-grade cloud automation and system connectivity. We are looking for a Senior Software Engineer to join our global Automation Engineering team to design, scale, and govern our enterprise integration substrate using AWS cloud services and modern iPaaS platforms. This is a senior individual contributor role for a hands-on system engineer who can set technical standards, establish integration best practices, and partner with cross-functional teams to automate complex business workflows at scale. What You'll do : Lead technical design and execution for automation initiatives within the team, creating paved paths that enable builders across Okta to connect enterprise systems seamlessly. Design, build, and deploy high-throughput event-driven integration flows , API gateways, and async orchestration workflows using AWS architectures and iPaaS platforms. Build reusable frameworks , developer SDKs, self-service primitives, and integration templates to streamline automation delivery. Serve as a technical domain expert on AWS cloud services and modern iPaaS tooling, driving scalable architecture, reliability, and builder enablement. Partner with operations, security, and platform teams to strengthen monitoring, observability, structured audit logging, and automated governance. Mentor team members and participate in cross-functional design reviews , instilling a platform engineering mindset and raising the b

typescriptpythonjava
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng

pythonsqlpostgresql
View job →
O
25 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Role Overview As the Okta Research & Design (ORD) Jira Owner & Administrator , nested within the Technical Program Management (TPM) organization, you will define and execute the strategy for how ORD plans, tracks, and reports on engineering work within the full toolchain of Atlassian Jira, Jira Advanced Roadmaps (Plans), Confluence, Rovo, and AI integrations. In this role, you act as a critical bridge between product & technical leadership, agile teams, and cross-functional business units — translating methodology, best practices, and process frameworks into practical workflow improvements that accelerate delivery and reduce manual overhead for a 1,000+ person engineering, product, and design organization. By optimizing tooling configurations, automating reporting, and eliminating fragmented management processes, you will directly accelerate ORD’s developer velocity and strengthen operational efficiency. Core Responsibilities 1. Platform Ownership & Strategic Tooling Direction Strategic Vision: Define, implement, govern an AI-led strategy for utilizing Jira for product planning, tracking, and reporting tooling that supports end-to-end portfolio tracking from high-level corporate objectives down to sprint-level stories. E2E Ecosystem Management: Complete ORD ownership of the toolchain (Jira, Jira advanced Roadmaps, AI integrations, Confluence), overseeing architecture, schema updates, workflow management, screen/notification schemes, cu

sqlawsci/cd
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description Okta is the leading independent provider of enterprise identity. The Okta Identity Cloud enables organizations to securely connect the right people to the right technologies at the right time. With over 6,500 pre-built integrations, Okta customers can easily and securely use the best technologies for their business. Over 7,950 organizations — including JetBlue, Nordstrom, Slack, and Twilio — trust Okta to protect the identities of their workforces and customers. Position Description We are looking for an experienced Senior Software Engineer – UI to join our Identity Platform engineering team. You will own the design and delivery of complex, enterprise-grade frontend experiences that power Okta's identity lifecycle management capabilities — including admin configuration flows, wizard UIs, real-time progress tracking, and bulk operation workflows. You will partner closely with Product Management, UX, and backend engineers to translate complex enterprise identity requirements into intuitive, accessible, and performant web applications. This is a hybrid position. Job Duties and Responsibilities - Frontend Ownership: Independently own and deliver complex UI features end-to-end — from design collaboration through production deployment. - Architecture & Standards: Lead frontend architectural decisions, enforce code quality, accessibility, perfor

javascripttypescriptreact
View job →

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous

pythondockerkubernetes
View job →

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →
G
Godaddy
📍 British Columbia• Full-time• From C$107K/yr
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operati

pythonawsci/cd
View job →

Location Details: At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operat

pythonawsci/cd
View job →
🔔

Get new senior infrastructure architect jobs by email

Daily job updates · Unsubscribe anytime