Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Senior Security Operations Engineer you will: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Harden our cloud-native environments (AWS, OCI, GCP) by introducing secure by default designs and features into network, tooling, and processes Own and drive resolutions for enabling engineers to design, build, and use infrastructure securely at scale by deploying secure architectures using infrastructure-as-code and reusable code libraries Manage IAM / RBAC for cloud infrastructure, and partner with IT on streamling authentication/authorization to ensure unified access control across the board Deploy and operationalize some of the security services and tools (eg: SIEM, SOAR, domain monitoring, endpoint tooling, cloud security tooling) Respond to security incidents and harden environments post-incidents. Support control monitoring and remediation for compliance initiatives Gather and analyze security metrics to address security issues with cross-team dependencies Be a problem solver who is empathetic to developer concerns and will employ construc
Jobiba hiring network
Senior Infrastructure Architect Jobs
7,292 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior infrastructure architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated
Senior Product Architect, K8s-Based AI Infrastructure — 2 Locations. Apply via Workday.
About Vercel: Vercel is the agentic infrastructure company. We free people and agents to ship what’s next. For more than a decade, Vercel has shaped how the web is built. As the team behind Next.js, v0, and AI SDK, we create products that help builders move from idea to production with speed, security, and exceptional developer experience. Now, software is entering a new era, and the next generation of products will not just be used by people. They will be built, extended, and operated by agents. We are building the platform for that future, trusted by companies like OpenAI, PayPal, Ramp, Supreme, and millions of developers worldwide . Whether you’re building our products, supporting our customers, growing our community, or shaping our story, you’ll help define what comes next. About the role: We are looking for a Senior Manager of Solutions Architecture to join and lead our our APAC SA team. In this role, you will be responsible for leading a team of Solutions Architects who support our sales teams across pre and post sales. You will drive technical excellence, mentor the team, and partner closely with sales leadership to win new business and expand existing accounts. This position will be based in Sydney, Australia. What you will do: Lead, develop, and manage an APAC SA team, including training and career development. Ensure technical excellence across the pre and post sales cycle. Provide hands-on support for key accounts as needed (presentations, architecture reviews, migration plans). Oversee and coach SA engagement on strategic accounts and critical deals. Partner with our engineering org to provide product feedback from the field. Develop repeatable plays from trends your are seeing across your team. Track and report SA performance to Sales and Field Engineering leadership quarterly. Own org planning, staffing, budgeting, and regional recruiting. About you: 8+ years of experience in Solutions Architecture/Engineering, Sales Engineering, or Technic
NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Role We are seeking a Pre-sales Architect to join our team of identity security experts. In this role, you will leverage your technical expertise and business acumen to drive customer success, helping organisations modernise and secure their identity ecosystems. As a trusted advisor, you will collaborate with enterprise customers, channel partners, and internal teams to deliver innovative identity and access management (IAM) solutions that align with business goals. This role requires a balance of strategic thinking, technical depth, and exceptional communication skills to translate complex identity challenges into solutions that address customer needs. You’ll play a pivotal role in accelerating sales cycles, enabling technical wins, and shaping the future of identity in your region. The Opportunity Reporting into the EMEA Office of the Field CTO, this role will be challenged daily to show prospects and customers how Okta can help solve their business challenges. We are searching for individuals that have experience deploying identity and access management solutions in large and complex environments. This hands-on experience and exposure to the operational aspect of real customer deployments is what sets you apart and provides credibility to the architectures you design and propose. Your proven experience with other products and legacy solutions will be invaluable when assessing a customer’s existing challenges and planning the
Datadog’s Implementation Services team helps customers implement and deploy Datadog quickly and successfully. Our team of architects leads the discovery, design, build, and launch of the Datadog platform to help customers accelerate time to value and get the most out of their investment. As a Senior Services Architect focused on Security and Cloud SIEM, you will help customers design, implement, and operationalize Datadog’s security capabilities across cloud, infrastructure, application, and log data sources. You will lead structured, outcome-driven professional services engagements delivered through a day-based professional services delivery model, partnering directly with customers through co-development working sessions, architecture workshops, implementation planning, and operational handoff. This role is ideal for someone who combines customer-facing consulting experience with strong cybersecurity knowledge, hands-on Cloud SIEM implementation skills, and an understanding of security control frameworks such as NIST 800-53, the NIST Cybersecurity Framework, CIS Controls, MITRE ATT&CK, SOC 2, PCI, HIPAA, ISO 27001, or similar standards. At Datadog, we place value in our office culture — the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design and guide execution of Datadog implementations, focusing on Security and Cloud SIEM deployment, including discovery, requirements gathering, technical architecture, deployment planning, and launch. Partner with customers to map security requirements, controls, and monitoring objectives to Datadog capabilities, including frameworks such as NIST 800-53, NIST Cybersecurity Framework, CIS Controls, MITRE ATT&CK, SOC 2, PCI, HIPAA, ISO 27001, or similar standards. Advise customers on security data strategy, including log source prioritization, parsing, normalizat
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a Senior Technical Architect on our Advanced Services team, you will bridge pre-sales strategy and post-sales delivery to architect enterprise data migration, disaster recovery, and hybrid-cloud storage solutions. In this dual consultative and hands-on role, you will collaborate closely with Sales, Engineering, and enterprise customers to translate complex business needs into resilient, automated IT infrastructure. You will drive revenue growth and project success by serving as the trusted technical authority across high-impact data center transformations. WHAT YOU'LL DO Architect & Scope High-Impact Solutions: Partner with Sales, Product Marketing, and customer leadership to evaluate complex enterprise environments, design high-availability storage architectures (replication, clustering, cloud integration), and scope Professional Services engagements that drive sales conversion and solution adoption. Lead Enterprise Data Migrations & Implementations: Own end-to-end technical execution of complex data center consolidations and block/file migrations using specialized appliances (e.g., Cirrus Data) and native toolsets, ensuring seamless cutovers with zero data loss and minimal operational disruption. Automate & Modernize Storage Operations: Develop custom scripts (Python, PowerShell, Shell) and integration workflows to automate replication, disaster recovery, and cloud-bursting processes within client
WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: As a member of the Global Technical Operations (TechOps), you will be a part of a team that focuses on operational reliability within a cloud-based infrastructure. You have hands-on cloud experience in architecting, building, deploying, managing databases, compute instances, and storage buckets. You have a passion for providing solutions through automation. You know that success is through collaboration and communication. What you'll be doing: Work in cross-functional teams to develop solutions and identify opportunities to bring efficiency and effectiveness. Research, evaluate, and incorporate new technologies/concepts into existing frameworks. Proactively identify areas to improve efficiency and effectiveness, recommend and implement solutions towards them. Develop and innovate operational practices, procedures for workflows, and documentation. Implement and contribute to IT security best practices. Automate tasks to ensure consistency and speed of deployment. Identify, analyze, and troubleshoot issues and work towards resolution. Explain technical solutions to bo
About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
NVIDIA pioneers computer graphics, gaming, AI, and accelerated computing. We are looking for a Senior Solution Architect with full-stack software engineering experience to join our team and play an important role in developing Sales AI applications. This position offers the opportunity to design, build, and evolve solutions that bring generative AI and intelligent workflows into everyday sales experiences. You will work across the application stack and collaborate with Product, AI and machine learning, Data, Security, Solution Architecture, and Engineering teams to deliver secure, reliable, and scalable solutions used globally. What you’ll be doing: Collaborate with application teams to design, develop, and maintain scalable full-stack solutions for enterprise sales workflows. Guide technical solutions across front-end, back-end, APIs, data services, integrations, and cloud infrastructure. Translate product requirements and business needs into secure, maintainable solutions and intuitive user experiences. Integrate generative AI models, AI services, APIs, retrieval systems, and agentic workflows into production applications. Design application architectures that support performance, availability, observability, security, scalability, and long-term maintainability. Lead technical design discussions, compare implementation approaches, make informed architecture decisions, and evaluate emerging technologies. Improve engineering practices for testing, code quality, continuous integration and delivery, monitoring, documentation, and production readiness. Investigate complex issues and develop solutions that improve reliability and user experience. Mentor engineers, share technical knowledge, and contribute to engineering standards and collaborative team practices. What we need to see: <
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Core Infrastructure team's mission is to build and evolve the foundational platform that every Robinhood engineering team builds on — owning the systems, primitives, and developer-facing abstractions that power 24/7 trading, crypto, and global expansion. We treat infrastructure as a product: reliable, fast to provision, and invisible to the teams above it. As a Senior Staff Software Developer on Core Infrastructure, you will own the architectural evolution of three deeply interconnected domains: service mesh and connectivity, compute platform, and infrastructure provisioning. Your decisions will directly shape engineering velocity, operational reliability, and Robinhood's ability to expand to new regions and markets. This is not an operations role — it's a once-in-a-platform-lifecycle opportunity to redesign the foundation before complexity becomes permanent! This role is based in our Toronto, ON office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-p
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur
NVIDIA is seeking a world-class computer architect to contribute to the development of future high-performance computing systems, with a focus on enhancing the power-constrained performance of the hardware. Ideal candidates will have a strong track record of understanding and analyzing memory systems architecture to improve performance per watt (perf/W) and performance per millimeter (perf/mm). A broad perspective across the field of computer architecture and depth in the area of power, performance, and area (PPA) analysis is highly desirable. NVIDIA has pioneered programmable GPUs and the CUDA language and is a world leader in high-performance computing technology, with aggressive plans for future processors. This position offers the opportunity to have a real impact in a fast-moving, technology-focused company. What you will be doing: Develop innovative high-performance processor and system architectures, focusing on the memory system and energy efficiency. Develop architecture and micro-architecture features to improve the state-of-the-art in GPU memory systems, optimizing along the axes of perf/W, perf/mm, and perf/$. Develop and enhance architecture prototype models for power and noise analysis. Participate in performance and power simulation of features to analyze, define, and improve energy per byte. Analyze benchmarks, application workloads, and performance/power simulation and emulation results to identify areas for architecture optimizations. Debug power, performance, and functional issues with high-level models, RTL simulation and emulation, silicon, and systems. Collaborate with outside partners on system infrastructure. What we want to see: 10+ yrs of experience in CPU/GPU architecture, memory systems design with a focus on energy efficiency in the system. Bachelor
The Development Infrastructure team builds the tooling and systems our Asana engineers use every day to bring their ideas to production quickly and reliably. We build and operate the software that drives Asana’s roadmap. Each day, we combine industry best practices and innovation to support this product-focused company. We’re looking for an experienced Software Engineer with a passion for developer infrastructure. You will work with a world-class team of engineers on deploying and operating existing developer tooling, and building new tools to support our global, growing development team. You will have a unique opportunity to design and develop the systems and applications that drive the Asana development experience, lead complex technical projects, and work on cross-functional initiatives to help define the future of software engineering at Asana. This role is based in our Reykjavík office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Innovate the architecture of our development sandboxes to accelerate development iteration. Modernize our build system, tighten iteration loops, and polish frequent cycles in the developer experience. Lead complex developer infrastructure projects from technical design through implementation, rollout, and operation. Partner with engineering teams to identify opportunities to enable teams to develop faster at Asana and help the company achieve our goals faster. Analyze complex developer infrastructure systems to uncover issues, root causes, and areas for improvement. Keep Asana up to date on open-source trends and identify new opportunities for improved development infrastructure. Champion co
Get new senior infrastructure architect jobs by email
Daily job updates · Unsubscribe anytime