Jobs in United States

Systems Architect in United States

4,982 active opportunities · Updated October 2026

Explore current systems architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-service tailored to AI applications that can be deployed anywhere and scale without limitations. This service supports the whole NVIDIA critical business from graphics drivers to autonomous vehicles to deep learning frameworks. To achieve this goal, we are looking for an engineer with a deep understanding of distributed systems, outstanding design skills, and a track record in building and delivering large-scale distributed services. What you will be doing: Leading the overall architecture and design of our distributed storage service optimized for AI/ML Develop and maintain distributed, robust and scalable Go programs deployed to state of the art open-source ecosystems, including Kubernetes. Develop and maintain user-space applications, containers, Go-bindings, and CLI tools. Building features for a distributed storage service to enhance availability and reliability for large-scale deployments Engaging and collaborating with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver Cloud services. Automating distributed storage service end-to-end, including deployment, management, and monitoring What we need to see: Bachelor’s of Science in Computer Science, or related field (or equivalent experience) with 8+ years of industry experience Strong background in developing distributed systems involving Golang, Kubernetes, and Cloud Service Provider integrations Strong track record of delivering distributed services in a variety of distributed computing environments Experience in i

KubernetesArtificial IntelligenceAIGolang
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

NVIDIA is a leading artificial intelligence computing company, and we are paving the way with innovations in self-driving cars, machine learning, supercomputing, gaming, and visualization. We give automakers, tier-1 suppliers, automotive research institutions, and start-ups the power and flexibility to develop and deploy breakthrough artificial intelligence systems for self-driving vehicles. Our unified computing architecture enables training deep neural networks in the data center, and then seamlessly runs them on NVIDIA DRIVE Platforms inside the vehicle. The Hypervisor and RTOS Team within NVIDIA DRIVE Software plays a critical role in NVIDIA's expansion into the world of artificial intelligence and autonomous vehicles. Our job is to facilitate the sharing and separation of system resources while achieving real-time, safety, and security requirements. We develop Hypervisor and RTOS with a strong focus on automotive quality, safety and security needed for the real-time, highly available system level components of world-class Autonomous Vehicles. We are making extensive use of formal methods to automate our workflow and increase the quality of our SW. We are hiring now for the position of Senior System Software Engineer for Hypervisor and RTOS What you’ll be doing: Design and develop new features for RTOS and hypervisor software stack. Bring up and optimize RTOS and hypervisor stacks on new NVIDIA Tegra SoCs. Develop high-integrity software using best-in-class engineering, safety, and security practices. Debug complex system-level issues across hardware, firmware, RTOS, and virtualization layers. Lead team-wide technical initiatives by building alignment, coordinating execution, and driving them to completion. What we need to see: BS, MS in CS/CE/EE or a related engineering field or equivalent experience </

Machine LearningArtificial IntelligenceAI
A
📍 California, USA - Remote, United States· Remote
✓ Quality checkedCompany trend +50%

Job Requisition ID # 26WD97363 Position Overview We are seeking a Principal Software Engineer – Backend to join Autodesk’s Enterprise Data Management (EDM) organization within the COO-GET Engineering group. This is a senior individual contributor role operating at the Principal (P4) level , expected to drive technology direction for large, complex, and business-critical backend and distributed systems . This role is anchored in backend software engineering excellence : designing, building, and evolving scalable services, APIs, and event-driven systems that operate at enterprise scale. As a Principal Engineer, you will work with high autonomy and ambiguity , shape long-term architecture, and influence multiple teams and domains. Familiarity with data engineering concepts is valuable, but backend systems, service design, and distributed systems are the core competencies. You will function as a technical authority and force multiplier—guiding design decisions, setting standards, and ensuring Autodesk’s core data services are reliable, resilient, and evolvable over time. Responsibilities Provide principal-level technical

PythonAWSAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole — fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements — including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software — ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software — health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior — champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW iss

M
📍 Richardson, TX, United States
✓ Quality checkedCompany trend -74.1%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Our vision is to transform how the world uses information to enrich life for all. Join an inclusive team passionate about driving innovation for next-generation memory solutions powering AI/ML and high-performance computing systems. As a Director of HBM RTL Design and Integration, you will lead a team responsible for the design, integration, and delivery of next-generation HBM SoC logic die, with a strong focus on RTL development and IP integration. You will drive technical strategy and execution across multiple product generations, working closely with architecture, verification, physical design, firmware, and product engineering teams. Key Responsibilities Lead SoC RTL design and integration for HBM logic die, including subsystem partitioning, IP integration, and SoC-level design convergence. Drive translation of architectural and micro-architectural specifications into robust, high-quality RTL implementations across multiple teams. Oversee SoC integration aspects, including clocking, reset, power intent, configuration infrastructure, and system-level design correctness. Establish design methodologies and best practices to improve quality, reuse, and development efficiency across HBM programs. Partner with SoC Architecture, Verification, Physical Design, Firmware, and System teams to ensure successful end-to-end product execution. Work closely with Product Engineering, Test, Probe, Process Integration, Assembly, and Manufacturing to ensure robust, manufacturable HBM

PythonAIRecruitment
B
📍 Berkeley, United States
✓ Quality checkedCompany trend +121.6%

Tactical RF Sensors Subsystem Lead Engineer Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking a Tactical RF Sensors Subsystem Lead Engineer to be the lead for sensor subsystem in Berkeley, MO , supporting a proprietary Air Dominance Fixed Wing program . You will be joining a fast-paced program where you will be responsible for developing and integrating complex, highly integrated, radio frequency (RF) sensor systems. In this role, you will be joining a cross-functional team integrating next generation capabilities as part of a very exciting and critically important development effort. Your role will be to oversee and coordinate the work of a team of engineers throughout the various stages of the program development life cycle. You will have the opportunity to operate in an agile development environment to mature our products, work on next generation open-architecture model-based designs, and to collaborate across multiple disciplines. The day-to-day activities of this position will be a combination of traditional design engineering, coordinating the daily activity of the engineering team, overseeing and approving technical work, collaborating with cross-functional peers, and presenting technical challenges and solutions to leadership. Frequent interaction with program leadership, customers, and industry partners during critical design phases will be key elements of this position. This position will provide the candidate an opportunity to develop their engineering skills with potential growth into a leadership role. Position Responsibilities: Coordinate the daily activity of an engineering team using agile toolsets (JIRA, Confluence, etc.) Provide technical l

Recruitment
B
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 Herndon, Vatican City State (holy See), United States
✓ Quality checkedCompany trend +121.6%

Java Software Engineer - Developer (Experienced and Senior) Company: The Boeing Company The Boeing Company is currently seeking Java Software Engineers – Developer (Experienced and Senior) to support our Advanced Ground Architecture team located in Herndon, Virginia, Colorado Springs, Colorado, Mesa, Arizona, Seal Beach, California and El Segundo, California. This position will focus on supporting the Boeing Defense, Space & Security (BDS) Software Engineering organization. The Advanced Ground Architecture (AGA) software team is a dynamic group of software engineers creating the future of Ground support with the extensibility and adaptability to be used across ALL Boeing programs. The software team is executing this vision through modern software technologies (Java, ReactJS, python, CI/CD pipelines) and methodologies (Scaled Agile). AGA is looking for self-motivated high performers to execute the large scope of Java development needed for the program's vision. The ideal candidates will provide software engineering functions for the design, development, and maintenance of complex, multi-tiered application software systems used to support the command and control of space vehicles. The software engineers will work day-to-day with system and test engineers in order to implement, test, and document new features and improvements for both web services and applications supporting distributed computing solutions. Position Responsibilities: Designs, develops, tests, and maintains software in an Agile execution that meets industry, customer, safety, and regulation standards throughout the end-to-end lifecycle Reviews, analyzes, and translates customer requirements into initial design and softwa

PythonJavaReactAngular
H
📍 New York, NY, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community You have shipped AI products before. You understand the difference between a demo and a production system. You have strong opinions about evaluation frameworks because you have experienced the consequences of operating without them. You are at your best when you own architecture decisions while continuing to build and deliver critical code yourself. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations back to source documents, and route complex cases to human experts. The output of these systems supports healthcare decisions that impact real members. As a Lead AI Applied Engineer, you will provide technical leadership for AI-enabled products and platforms, define architectural direction, establish engineering standards, and personally design and build the most critical components of our systems. You will lead through both technical expertise and execution, helping the team deliver reliable, scalable, and auditable AI solutions in a highly regulated healthcare environment. Why Join Us Lead the architecture of production AI systems where LLMs are foundational to the product experience. Make key technical decisions regarding model selection, system boundaries, platform architecture, and build-versus-buy strategies. Own the highest-risk and highest-impact technical challenges involving reliability, explainability, and correctness. Influence engineering culture and establish standards that shape how the team builds and ships AI products. Work on systems operating at meaningful scale, processing millions of documents and supporting healthcare decisions across a large member population. Partner

JavaScriptTypeScriptPythonReact
H
📍 New York, NY, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community Every large organization is making critical decisions today about how it will leverage AI over the next decade. Few have leaders who can both define that vision and demonstrate its viability through hands-on engineering. This role requires both. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations to source documents, and route difficult cases to human experts. These capabilities support decisions that impact real healthcare outcomes for members. As a Principal AI Applied Engineer, you will define the technical strategy, architectural standards, and long-term vision for AI-enabled products across the organization. You will influence enterprise-wide decisions regarding AI platforms, model strategies, engineering standards, and technology investments while remaining deeply hands-on in prototyping, experimentation, architecture, and software development. This is the highest-level individual contributor role within the AI Applied Engineering organization. Success requires exceptional technical depth, organizational influence, strategic thinking, and the ability to translate emerging AI capabilities into scalable, reliable, and responsible production systems. Why Join Us Shape the long-term AI architecture and engineering direction for a large enterprise healthcare organization. Influence how AI-enabled products are designed, built, evaluated, deployed, and governed across multiple teams. Drive strategic decisions involving models, vendors, platforms, infrastructure, and shared capabilities. Prototype and validate emerging technologies before the organization invests at scale.</

JavaScriptTypeScriptPythonReact
M
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

We're looking for a Senior Mechanical Engineer for Midjourney Medical — from precision electromechanical assemblies up to the large-scale structures and mechanisms that hold everything together and make it move. This role spans the full range of scale: one week you might be refining a compact transducer mount, the next you're architecting a structural frame or designing the motion system that positions hardware within it. This is a hands-on role on a small cross-functional team. You'll take problems from whiteboard sketch to working hardware yourself — designing in CAD, prototyping in the shop, testing, iterating, and integrating with research teams, electrical and software engineers along the way. We move fast and expect you to drive your own work: identifying what needs to happen next, making sound engineering calls without waiting for permission, and shipping hardware that works. What you'll do Own parts of the mechanical systems end to end — structures, mechanisms, and electromechanical integration Design large-scale structures: frames, weldments, enclosures, and support systems, with attention to stiffness, weight, manufacturability, and serviceability Design mechanisms: linkages, motion stages, actuation systems for precise, reliable positioning and articulation Integrate transducers, electronics, cabling, and thermal management into large physical systems Perform engineering analysis (tolerance stack-ups, structural/FEA, mechanism kinematics) to de-risk designs before committing to hardware Build and test relentlessly — 3D printing, rapid prototyping, and hands-on fabrication are core to how you'll validate designs Select parts and materials for system-level integration, balancing performance, cost, and lead time Drive projects independently from concept through validation, and collaborate closely with a small cross-disciplinary team to hit project goals What we're looking for Bachelor's degree in Mechanical Engineering or a related field 5+ years of experien

AIGoExcelSEM
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the low-level device runtime that turns compiled programs into efficient, functional and performant execution on OpenAI’s custom AI accelerator. This software will schedule kernel launches, manage device memory and address spaces, coordinate synchronization, and expose reliable abstractions to higher-level runtimes and frameworks. You will work at the boundary of software and hardware, partnering with compiler, kernel, architecture, verification, and silicon teams to define interfaces and validate behavior. You will also use and improve event-based, cycle-accurate simulation to develop runtime capabilities before silicon is available, diagnose performance and correctness issues, and guide hardware-software co-design. In this role, you will: Design and implement the low-level device runtime for OpenAI custom silicon. Build kernel-launch scheduling, command submission, queueing, dependency tracking, and completion handling. Manage device memory spaces, allocation, virtual-to-physical mappings, data movement, and lifetime across concurrent workloads. Implement synchronization primitives, events, barriers, streams, and ordering guarantees that are correct and efficient. Define clean interfaces between the runtime, drivers, firmware, compiler-generated code, kernels, and higher-level execution systems. Use event-based, cycle-accurate simulators to develop, validate, debug, and performance-tune runtime behavior before and after silicon availability. Di

AWSRestAIC++
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new

PythonAWSRestAI
C
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -100%

From $218K/yr

Quick readStrong listing-quality and freshness signals

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Staff Software Engineer on the Core Automation team within the Platform group, you'll architect and build the Agentic AI systems that are transforming how Coinbase operates. This team is reimagining customer support and compliance processes for a fully AI-driven world, designing intelligent agents, orchestration frameworks, and measurement systems that deliver delightful customer experiences at scale. You'll own the technical direction for production AI systems, working across cross-functional teams to bring this vision to reality while building primitives that scale automation across the company. What you'll do: Architect and build Agentic AI systems that power Coinbase's compliance automation and other Operations, from intelligent agents through orchestration and guardrails Design foundational APIs and measurement frameworks that ensure AI agents are grounded, relevant, and reliably deliver customer delight with minimal hallucination Lead technical direction for distributed systems underpinning AI automation, defining architecture patterns and strategic roadmaps in partnership with engineering leadership Build reusable primitives and orchestration solutions that enable AI-powered automation to scale across multiple domains beyond the initial customer support and compliance focus Mentor engineers on AI system design techniques, coding standards, and production-

AWSAIGoFinance
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as &#34;the AI computing company.&#34; We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa

PythonGitLinuxArtificial Intelligence
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each brings together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will be instrumental in ensuring our data center solutions deliver industry-leading performance for accelerated computing workloads. What you will be doing: Design and execute comprehensive performance benchmarking strategies for our data center platforms and products Characterize real-world AI training, inference, and HPC workloads at scale Define, track, and report key performance indicators (throughput, latency, efficiency, scaling) Build automation tools and frameworks for performance monitoring and analysis Identify and analyze performance bottlenecks across compute, memory, network and storage subsystems Work closely with architecture, hardware,

PythonDockerKubernetesLinux
🔔

Get new systems architect jobs in United States by email

Daily job updates · Unsubscribe anytime