NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated
Jobiba hiring network
Senior System Architect Jobs
15 active opportunities · Updated for September 2026
Fresh results
15 shown
Explore current senior system architect jobs. Use filters to narrow by work mode, employment type, experience and date posted.
NVIDIA is now looking for a Senior Memory System Engineer to join our ASIC Memory Subsystem team! As a Senior Systems Engineer at NVIDIA, you'll join a group of hardworking engineers to develop and architect innovative Memory Solution for Tegra SoCs. In this position, you'll make a real impact in a multifaceted, technology-focused company. You will work with memory controller/PHY and Platform / System architect, Firmware, SI/PI, Memory suppliers to design and architect cutting edge, high speed and lower power memory technology for NVIDIA CPUs and SOCs. What You Will Be Doing: Analyze future DDR/LPDDR/HBM technologies to determine optimum performance, power, function and RAS in memory for Next generation SOC and Systems. Collaborate with ASIC Architects, Designers, Software and Firmware SW/FW teams to drive memory technology and associated requirements for memory controllers. Define Memory module, Package, and PCB layouts appropriate to the system workloads Debug and bring up memory evaluation / validation and failure issues on memory technology. Collaborate with DRAM suppliers and industry partners on to develop memory and memory related component technology. What We need to see: Bachelor's degree or master’s degree in Electrical Engineering, Computer Engineering (CE), or a related field (or equivalent experience) 10 years of proven track record in DRAM design, module design, or memory sub system design. Deep understanding and strong fundamental of memory design, features, ECC algorithm, SI and PI (Training algorithm) in DDR, LPDDR, and HBM. Strong understanding of memory sub system level interaction with Cache, Memory controller and PHY. Experience in the design, bring-up and validation for memory failure analysis Experience with Python, C/C
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ’s System Software team enables on-device experiences to take full advantage of our hardware capabilities. We collaborate with internal and external partners in a high-leverage environment that enables us to spend a majority of our time developing solutions that are unique to our hardware, sensors, algorithms, and interaction models. If you enjoy solving complex, interdisciplinary problems with a world-class team, we'd love to hear from you! What You Might Do Learn what it's like to be a part of a world-class embedded software team building a first-of-its-kind product in a startup environment Responsible for system design and architecture Develop low-level driver and framework software in C and C++ Develop device-focused infrastructure software in Python Debug issues at the interface between hardware and software Optimize software for better performance and lower power consumption Collaborate in the software engineering process with documentation, testing, and code review Essential Qualifications 8+ years of experience in system software engineering and embedded platf
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. About the Role We are looking for Staff System Software Engineer in Test to join our team. In this role, you will be responsible for design, development, automation and reporting of Integration and system tests spanning across firmware and device drivers. This role requires you to have significant technical breadth and deep understanding of low-level system software specifically in server class systems. You will be part of a new team responsible for integration of different system software deliverables and development of system tests spanning all the components. You will contribute to shaping the test strategy , guide best practices and solve complex problems while maintaining a strong hands-on focus. You will partner with development and other QA teams to deliver high quality scalable and reliable solutions. About the Team Integration and system test team is responsible for verification and validation of integrated components across Board management controller (BMC), Firmware and Linux device driver. The team is also responsible for management and maintenance of common tools and pipel
NVIDIA's Deep Learning GPUs have ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are growing our company and the team with the smartest people in the world. We are looking for extraordinary Software Engineers to develop and productize NVIDIA's DRIVE OS software. As a member of NVIDIA's Solution Engineering team, you will adapt DRIVE OS solutions to various car platforms equipped with different sensors. We are looking to hire Senior System Software Engineer – AUTOSAR. Ideal candidate will have very strong programming skills, a good grasp of HW & SW Architectures, a solid exposure to AUTOSAR & related architecture, tools and frameworks. What you will be doing: Participate and provide inputs and recommendation into AUTOSAR Architecture evolution with design choices, tools and methodology Architectural explorations on both SW and HW fronts which include feasibility studies, quick prototyping, profiling, safety studies, data analysis and presentation of results Influence next-gen HW architectures and SW Architecture and design Drive complex technical issues to closure that may occur interacting with cross-teams What we need to see: BS/MS, or equivalent experience 5+ years of experience Strong programming skills in C/C++ and scripting skills in Perl, Python etc Good experience and com
NVIDIA is seeking a world-class computer architect to contribute to the development of future high-performance computing systems, with a focus on enhancing the power-constrained performance of the hardware. Ideal candidates will have a strong track record of understanding and analyzing memory systems architecture to improve performance per watt (perf/W) and performance per millimeter (perf/mm). A broad perspective across the field of computer architecture and depth in the area of power, performance, and area (PPA) analysis is highly desirable. NVIDIA has pioneered programmable GPUs and the CUDA language and is a world leader in high-performance computing technology, with aggressive plans for future processors. This position offers the opportunity to have a real impact in a fast-moving, technology-focused company. What you will be doing: Develop innovative high-performance processor and system architectures, focusing on the memory system and energy efficiency. Develop architecture and micro-architecture features to improve the state-of-the-art in GPU memory systems, optimizing along the axes of perf/W, perf/mm, and perf/$. Develop and enhance architecture prototype models for power and noise analysis. Participate in performance and power simulation of features to analyze, define, and improve energy per byte. Analyze benchmarks, application workloads, and performance/power simulation and emulation results to identify areas for architecture optimizations. Debug power, performance, and functional issues with high-level models, RTL simulation and emulation, silicon, and systems. Collaborate with outside partners on system infrastructure. What we want to see: 10+ yrs of experience in CPU/GPU architecture, memory systems design with a focus on energy efficiency in the system. Bachelor
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
EMPLOYER IS A CONTRACTOR FOR THE U.S. GOVERNMENT. THIS POSITION REQUIRES U.S. CITIZENSHIP. Role Description: We are seeking a talented and experienced Platform Engineer to join our team. Possessing strong proficiency with Kubernetes, the ideal candidate will play a crucial role in the development and maintenance of our platform, demonstrating a strong understanding of system architecture, problem-solving skills, and the ability to guide junior engineers. As a Platform Engineer within our organization, you will join a team of dedicated unicorns who are focused on advancing freedom and independence globally. Our teams work on various different technical baselines depending on specific engagements. We focus our Products on the Platform layer, but to deliver for our Heroes, we also work IaC automation, security, and implementation. Responsibilities: You will help us automate and build a runtime Production Environment and serve as a link between the Applications, either custom development, COTS, or OSS, and that environment, advocating best practices and facilitating communication between the application teams and platform teams. In this position, you will be expected to: Validate Solutions/Implementations: Ensure that solutions and implementations align with the outlined tasks and business requirements. Conduct thorough validations to maintain the integrity and efficiency of the platform. Independently Problem Solve: Demonstrate the ability to identify and solve business problems independently. Develop small components to address specific challenges without relying on explicit architecture diagrams. Understand Medium-Level Tasking: Comprehend how medium-level tasks contribute to the achievement of overall goals. Collaborate with cross-functional teams to integrate various components into a cohesive and functional platform. Estimate Time and Effort: Provide accurate time and effort estimates for medium-sized tasks. Assist in project planning and resourc
1671 About the Role We are seeking a Senior Signal Integrity Engineer to develop, validate, and optimize high-speed signaling solutions across blade- and rack-level architectures for advanced compute platforms. This role sits at the intersection of silicon, package, interconnect, board, and system design , with a strong emphasis on hands-on measurement, simulation correlation, and cross-functional technical communication. Key Responsibilities End-to-end signal integrity analysis for blade- and rack-level system architectures. Analyze and optimize high-speed and low-speed I/O interfaces, including PCIe Gen4/5/6, Ethernet, DDR, SerDes, SPI , I2C, etc . and related interconnects. Perform time-domain and frequency-domain simulations using tools such as Ansys HFSS, Keysight ADS, Cadence Sigrity, CST, SPICE , or similar. Support hands-on lab validation using VNA, TDR, BERT, and high-speed oscilloscopes . Correlate simulation results with lab measurements to identify margin gaps, debug issues, and improve design methodology. Collaborate with silicon, package, board, connector, cable, and system design teams to optimize I/O channel performance. Work with interconnect vendors and ODMs to guide board layout, stack-ups, routing rules, and system design decisions. Review schematics, layouts, simulation results, and validation data for blade, backplane, and rack-level hardware. Prepare and communicate clear validation reports, measurement summaries, debug findings, and technical recommendations to internal teams, vendors, and senior technical stakeholders. Required Qualifications Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, or a related field . Strong experience in signal integrity for high-speed digital systems. Hands-on measurement expertise using VNAs, TDRs, BERTs, and high-speed oscilloscopes . Experience measuring and analyzing S-parameters, impedance profiles, eye diagrams, jitter, timing margins, insertion loss, return loss, and cro
Join the NVIDIA's Solutions Engineering team that is reshaping the future of driving! Our goal is to build and deploy scalable solutions for autonomous vehicles and as a result, create safer and more efficient roads. Our team is hands-on, passionate about practical results, and values diversity. You will help craft the application software architecture by working closely with external partners developing on our platform and on collaborations across multiple teams within NVIDIA working on autonomous vehicles. You will also advance and refine the overall drivability of our solution, focusing on integration challenges and using your deep analytical skills to tease through the complexity of the system to find effective solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers, and is committed to fostering a diverse work environment and proud to be an equal opportunity employer. If you are passionate in bringing autonomous vehicles into the world and see the solution come together, we would like to hear from you! What you'll be doing: Shape the application architecture internally, with a focus on perception & sensor fusion, by collaborating closely with architecture and software development teams. Integrate and adapt NVIDIA solutions in target vehicles, ensuring that both perception and sensor fusion are adapted and tuned to meet the desired driving performance and functionality. Lead bring-up activities and provide technical support to resolve functional and perception & sensor fusion related issues. Perform and leverage in-vehicle and simulation test drives for functional and performance analysis on the recorded data. Work with our partners to efficiently integrate hardware and software components, understand the system architecture, profile performance, identify bottlenecks, and drive optimization <
Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Drives AI-first tooling, SDLC automation, and developer productivity initiatives Required Experience & Skills 6+ years of experience building internal tools or developer platforms Strong CI/CD and automation expertise, Generative AI integration in SDLC Core Backend & Architectural Fundamentals: Deep dive into data structures, algorithms, and microservices/distributed system architecture patterns. While we embrace a language-agnostic mindset, strong proficiency in Java or Python is preferred. Cloud & Infrastructure: Hands-on experience with cloud infrastructure, container frameworks (Docker/Kubernetes), and CI/CD pipelines is a major advantage Modern Engineering Practices (Highly Valued): Daily utilization of AI tools and agentic engineering environments (e.g., Kiro, Claude, Copilot) to accelerate software design, automate background workflows, and optimize developer productivity Tools and Resources Training and Development: During onboarding, we focus on equipping new hires with the skills and knowledge for success through department-specific training. Continuous learning is a central focus at Appian, with dedicated mentorship and the First-Friend program being widely utilized resources for new hires. Growth Opportunities: Appian provides a diverse array of growth and development opportunities, including our leadership program tailored for new and aspiring managers, a comprehensive library of specialized department training through Appian University, skills based training, and tuition
Role: Senior AI Engineer Location: Hyderabad, India (Hybrid) Department: Product Development About the Role GHX is building a cutting-edge LLM-powered document understanding platform focused on classification, structured data extraction, and intelligent orchestration at scale. This is a high-impact AI engineering role where you will own the full lifecycle—from problem framing to production deployment . Initially, you will focus on prompt engineering and evaluation systems , building the quality foundation for AI performance. Over time, the role expands into agent orchestration, system architecture, and migration of rule-based systems to LLM-driven pipelines . A strong foundation in software engineering (5+ years) is essential. This role demands engineering rigor across both traditional system design and AI system behavior . Core Responsibilities 1. Prompt Engineering Design prompts for diverse document classification and extraction tasks Treat prompts as formal specifications (precise, structured, and edge-case-aware) Develop few-shot, chain-of-thought, and structured output templates Manage prompt lifecycle: versioning, testing, and rollback 2. LLM Output Evaluation Create and maintain ground truth datasets Build automated evaluation pipelines (precision, recall, field-level accuracy) Identify and resolve conceptually incorrect outputs despite surface correctness 3. AI Agent Orchestration Design multi-agent workflows for document processing Implement tool-use patterns and integrate MCP servers Optimize orchestration for scale and efficiency 4. Software Engineering Develop production-grade APIs and backend services Apply Clean Architecture / DDD principles Write maintainable, testable Python code Contribute to CI/CD, deployment, and observability systems 5. Stakeholder Collaboration Act as a bridge between business stakeholders and AI systems Translate product requirements into technical architectures Communicate system behavior, limitations, and quality
We are hiring Senior / Lead Engineers to design, build, and scale robust, production-grade data and platform systems for a product portfolio within Blenheim. This is a hands-on engineering role where you will own complex problem areas end-to-end, influence system architecture, mentor engineers, and deliver scalable solutions in a fast-growing, product-led environment. If you enjoy working with ambiguity, building platforms from the ground up, and shipping real business impact, we’d love to hear from you. About Blenheim Chalcot: Blenheim Chalcot India is part of Blenheim Chalcot, a global venture builder headquartered in London. With over 26 years of innovation, we've been at the forefront of creating some of the most groundbreaking GenAI-enabled companies. Our ventures lead the charge in digital disruption across a spectrum of industries, from FinTech to EdTech, GovTech to Media, and beyond. Our global presence spans the US, Europe, and Southeast Asia, with a portfolio that employs over 3,000 individuals, manages assets exceeding £1.8 billion, and boasts total portfolio sales of over £500 million. The Role: As a Senior / Lead Engineer, you will play a key role in designing and delivering scalable systems across core platforms. You will collaborate closely with Product, Data, and Engineering teams in Mumbai and London, helping shape technical direction while remaining hands-on. Key Responsibilities: Successful candidates will take a leading role within our Engineering Centre of Excellence in Mumbai, with day to day responsibilities which will include. Lead and mentor a team of engineers, fostering a culture of collaboration and continuous improvement. Oversee the planning and execution of projects, ensuring alignment with business objectives and timelines. Provide technical guidance and expertise to the team, promoting best practices in engineering. Collaborate with cross-functional teams to understand requirements and deliver effect
Job Title: Senior QA Engineer - Performance Testing Paytm is India's leading mobile payments and financial services distribution company. A pioneer of the mobile QR payments revolution in India, Paytm builds technologies that empower small businesses with payments and commerce solutions. Paytm’s mission is to serve half a billion Indians and bring them into the mainstream economy through the power of technology. About the Role: We are seeking a skilled Performance Test Engineer to design, execute, and analyze performance tests to ensure application scalability, stability, and responsiveness under varying load conditions. The ideal candidate will have hands-on experience with industry-standard performance testing tools and a strong understanding of system architecture, monitoring, and troubleshooting. Expectations/ Requirements Develop comprehensive performance test strategies and plans aligned with system requirements, project timelines, and business goals. Understand application architecture and identify critical business transactions for performance validation. Design realistic workload models to simulate real-world usage scenarios. Create, maintain, and execute performance test scripts using tools such as JMeter, LoadRunner, Gatling, or similar. Conduct baseline, load, stress, and scalability testing to evaluate system behavior under different conditions. Monitor system performance using tools like Influx DB, Grafana, JVM monitoring tools, and MAT (Memory Analyzer Tool). Analyze test results to identify performance bottlenecks and system limitations. Collaborate with development and infrastructure teams to troubleshoot and resolve performance issues. Assess system scalability and recommend optimizations to improve performance and reliability. Generate detailed performance test reports, including metrics, findings, and actionable recommendations. Work with stakeholders to gather and validate Non-Functional Requirements (NFRs), SLAs, and KPIs. Perform API an
Get new senior system architect jobs by email
Daily job updates · Unsubscribe anytime