NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-service tailored to AI applications that can be deployed anywhere and scale without limitations. This service supports the whole NVIDIA critical business from graphics drivers to autonomous vehicles to deep learning frameworks. To achieve this goal, we are looking for an engineer with a deep understanding of distributed systems, outstanding design skills, and a track record in building and delivering large-scale distributed services. What you will be doing: Leading the overall architecture and design of our distributed storage service optimized for AI/ML Develop and maintain distributed, robust and scalable Go programs deployed to state of the art open-source ecosystems, including Kubernetes. Develop and maintain user-space applications, containers, Go-bindings, and CLI tools. Building features for a distributed storage service to enhance availability and reliability for large-scale deployments Engaging and collaborating with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver Cloud services. Automating distributed storage service end-to-end, including deployment, management, and monitoring What we need to see: Bachelor’s of Science in Computer Science, or related field (or equivalent experience) with 8+ years of industry experience Strong background in developing distributed systems involving Golang, Kubernetes, and Cloud Service Provider integrations Strong track record of delivering distributed services in a variety of distributed computing environments Experience in i
Jobs in United States
Senior System Architect in United States
1,941 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior system architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each brings together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will be instrumental in ensuring our data center solutions deliver industry-leading performance for accelerated computing workloads. What you will be doing: Design and execute comprehensive performance benchmarking strategies for our data center platforms and products Characterize real-world AI training, inference, and HPC workloads at scale Define, track, and report key performance indicators (throughput, latency, efficiency, scaling) Build automation tools and frameworks for performance monitoring and analysis Identify and analyze performance bottlenecks across compute, memory, network and storage subsystems Work closely with architecture, hardware,
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We’re searching for a highly motivated, technical leader to design, drive, and operationalize rack-scale factory and deployment flows for next-generation data center products. The ideal candidate will combine deep systems expertise, decisive technical leadership, and a passion for building reliable, debuggable, and scalable manufacturing and deployment solutions. What you’ll be doing: Lead and drive rack-scale/L11 flows for factory and initial data center deployment. Design and implement end-to-end factory workflows, including firmware flashing sequences, security provisioning, and deployment of software mitigations. Collaborate with data center architects, ODMs, and OEMs to define factory and data center requirements that ensure efficient and reliable production ramp. Champion reliability, debuggability an
About Us: Blockworks is an information platform that sits at the center of the crypto industry. We transform raw, complex data and facts into actionable research, trusted alpha-driven insights, and world-class events. The result is transparency and confidence. Blockworks connects investors and businesses in onchain capital markets. We give businesses a platform to earn trust and provide investors with the information they need to underwrite the asset class. Who You Are: You have a keen focus on backend and API development and Software engineering is your passion. You are a player-coach and a natural leader who understands the technical and human elements that go into great software design. You have a results-oriented attitude and a passion for delivering flawless releases and developing digital product pipelines (CI/CD pipelines). You have a proven track record facilitating engineering teams to increase productivity and quality. You're excited at the possibility of being on the ground floor of the design and development of backend strategies. You bring a passion for designing and maintaining scalable API services that handle large amounts of data elegantly. You love moving quickly in a fast-paced start-up, but you also bring intentionality, sustainability and scalability to your approach as an engineer. What You’ll Do: As a Senior Backend Engineer at Blockworks, you’ll design, build, and maintain the systems that power our products end-to-end. From high-performance APIs to database architecture, you’ll own the backend layer that makes everything else possible. You won’t just be handed requirements, you’ll help define them, scope projects, and make the architectural decisions that shape our technical foundation. Your work will directly impact our research platform ( blockworksresearch.com ) and our media site ( blockworks.co ; 1M+ monthly active users). Every day will look a little different, but in general, you will do things like: Architect, build, and ship backend
From $184K/yr
TPMs at Datadog see the problems hiding between teams, engineer away the work that shouldn’t require humans, and drive the company’s most technically complex and consequential bets to completion. Technical Program Management at Datadog operates at the intersection of engineering depth and organizational reach by driving high priority, cross-functional programs that are too complex and consequential for any single team to own. We partner with engineering on solving deeply technical problems at scale by connecting the people, decisions, and context to move Datadog's most important work forward. We build the systems and automation that make entire classes of program work self-executing. We are in the architecture conversation early, earning trust through technical judgment. We use AI to surface risks earlier, accelerate program execution plans, and find cross-team patterns that would otherwise stay hidden. The faster teams move, the more essential it is to have someone who can operate across them. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What We Expect: These are the expectations we hold for every TPM at Datadog. Technical depth, product domain expertise, and AI systems literacy; knowing how AI solutions work, where they fail, and the scope and impact of those failures. AI brings more complexity into the picture - the technical bar is higher, not lower. Build the systems that reduce the need for coordination Identify what matters before anyone asks, and automate the rest Engineer program lifecycles end-to-end See what no single team can see and own the solution Drive the company's most technically complex and consequential bets through cross-functional agreement, organizational visibility, and influence Build AI powered automation tools and
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a skilled and experienced Senior AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay curre
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's Cache team is building a next-generation caching solution designed to deliver sub-millisecond average latency, horizontal scalability, and high efficiency—all at a drastically lower cost. Our ultimate vision is to shape a caching infrastructure capable of supporting 1 billion Daily Active Users while reducing costs by 90%. We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to solve Roblox's ever-growing caching challenges. You will report directly to the Engineering Manager for the Cache team. (Check out our recent engineering blog post here to learn more about the team's latest work!) You will: Lead the architectural transition to a next-generation, multitenant caching service built on ValKey, ensuring strict data, resource, and failure isolation for all tenants. Drive systemic optimizations to mitigate head-of-line blocking, manage hot keys, and maximize CPU and memory utilization across physical machine clusters. Design and build robust frameworks to a
At Datadog, we’re on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Observability Data Platform (ODP) is the backbone of everything Datadog delivers – powering how data is ingested, stored, routed, and surfaced across every product at planet scale. As a Senior Product Manager for ODP, you will work with world-class engineers and cross-functional partners to shape how the platform is deployed, controlled, and operated. You will define product direction across the control plane and data layer, translate complex infrastructure trade-offs into clear roadmap decisions, and help customers get the most from their observability investment – regardless of architecture, topology, or scale. At Datadog, we place value in our office culture – the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You Will Do: Develop a deep understanding of the Observability Data Platform customers – platform engineers, SREs, and product managers that own the product verticals – their infrastructure challenges, deployment topologies, and cost-to-serve trade-offs. Define product direction across multiple ODP surfaces, including the control plane and data layer, by articulating clear problem statements and desired outcomes, and partnering with engineering on technical approach and sequencing Lead conversations with design partners and strategic customers to understand real-world platform pain points, validate product assumptions, and guide solutions from early prototypes through General Availability Develop a co
From $92K/yr
We are hiring a Senior Technical Product Marketing Manager to lead positioning and messaging and to grow adoption of MongoDB Search and Vector Search as foundational components of our platform – the retrieval layer powering the next generation of grounded AI applications and agents. This is a high-impact role for a marketer who thinks like a builder. As developers architect increasingly sophisticated systems – RAG pipelines, agentic workflows, multi-modal search experiences – retrieval has moved from an implementation detail to a core design decision. You’ll join a high-performing, globally distributed team and partner closely with Marketing, Builder Relations, Product Management, Engineering, Partners, and Sales to develop, measure, and achieve cross-functional goals. The role requires technical depth in information retrieval — lexical and vector search, hybrid approaches, embeddings, re-ranking, agentic retrieval loops, and the tradeoffs that matter in production systems — paired with the product marketing instincts to turn that depth into crisp, differentiated messaging for distinct user and buyer personas. Hands-on experience building or shipping AI-enabled products is a strong advantage. Individuals with prior experience in technical sales, developer relations, or technical marketing are encouraged to apply. This person is a voracious consumer of AI research and pays close attention to shifting patterns in application architectures and development, including agentic systems. This individual is confident in communicating with technical practitioners and non-technical decision makers in one-to-few and one-to-many engagements for internal and external audiences. We are looking to speak to candidates who are based in the US for our hybrid working model. What You’ll Do Drive Strategy & Execution: Act as a strategic partner for high-impact initiatives that align with MongoDB’s long-term business goals in collaboration with Marketing, Developer Relations, Product
NVIDIA is a global leader in high-speed computer vision, artificial intelligence (AI), and deep learning. Our team develops data engineering solutions that empower AI developers in autonomous vehicle (AV) domains to innovate quickly and effectively at scale. Are you ready to take on a senior technical role in building high-performance AI data pipelines? We seek an exceptional individual to design and optimize microservices and data pipelines to process massive volumes of AV data and enable seamless data mining and AI training. The ideal candidate will bring expertise in big data processing and distributed computing to create efficient solutions and overarching architectures for challenges such as video data curation, behavioral search, and AI dataset management. What you'll be doing: Scope and build tools, microservices, workflows, and distributed applications to accelerate data mining and AI training. Design and implement solutions for streaming, resilience, logging, security, authentication, workflow orchestration, and data management. Deploy AI models. Design and develop Retrieval-Augmented Generation (RAG) workflows enabling hybrid and agentic patterns. Analyze and operationalize complex distributed systems for speed-of-light performance. What we need to see: Experience developing high-performance, scalable software systems. MS with 6+ years, or BS (or equivalent experience) with 8+ years of relevant experience in Computer Science, Computer Engineering, or a related technical field. Strong programming skills in Python or Golang Proficiency in key technologies like Kubernetes, Helm, Hive, Parquet, SQL, vector databases, e.g., Milvus. Strong architectural skills with a proactive, problem-solving mentality. Experience in data mi
Become a part of our caring community The Senior Full Stack Engineer Performs software engineering activities in all layers of the stack, from setting up the database to programming in the back-end and the appearance at the front-end. The Senior Full Stack Engineer works on problems of diverse scope and complexity ranging from moderate to substantial. You will report to the Associate Director of Software Engineering. The Senior Full Stack Engineer is involved in all stages of software development. This includes front-end development, back-end development, database integrations, network and hosting management, user interface, user experience, and back-end server management. You are experienced with Cloud-based services and development end-to-end. You can work with autonomy while collaborating with others on architectural and process decisions. You are comfortable with giving and receiving feedback from your teammates and others in the organization. You conceptualize what you need to do from incomplete and evolving requirements. Experience communicating updates and resolutions to customers and other partners. On any given day you may: Contribute to technical architecture reviews and design as well as present vision Expertise using advanced data structures and algorithms Possess the ability to evaluate problems in both strategic and tactical terms Mentor and assist in coaching a small team of high-caliber engineers Drive positive change with technical vision and innovative solutions Strong affinity to software-driven engineering & automation Partner with the business and stakeholders to define technical direction & work cross functionally Deliver business-critical systems and compon
$225K – $300K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Software Engineer, Data, you will design, build, and operate the next generation of our data platform and products – going beyond ID to power a networked digital identity – while keeping member privacy, security, and reliability at the core. What you’ll do: Build and operate scalable, reliable data systems and pipelines – from ingestion to modeling to visualization – so Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner. Develop and maintain end-to-end data products and pipelines (batch and/or streaming) that collect, clean, transform, and model data, and own the infrastructure that powers them to unlock new business use cases and reporting. Implement and maintain infrastructure-as-code, CI/CD, and shared developer tooling for data products (e.g., Pulumi/Terraform, GitHub, orchestration tools like Dagster/Airflow) to make it easy and safe for teams to build, test, and ship changes across environments. Improve the security, compliance, and cost posture of the data stack through robust dependency management, IAM and secrets hardening, observability, and performance/cost optimizations. Partner with product and other stakeholders to uncover requirements, make architectural decisions, and continuously improve our data platform and processes. How you’ll measure success: Data reliability & SLAs: % successful pipeline runs, adherence to freshness SLAs for core datasets, and reduction in data-related incidents impacting stakeholders. Platform quality & efficiency: Reductio
$225K – $300K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. Today, CLEAR is well-known as a leader in digital and biometric identification, reducing friction for our members wherever an ID check is needed. We’re looking for a Senior Software Engineer to establish our Observability framework and foundations. You will join us to accelerate building and scaling our innovative systems that support our growing identity platform. You will drive on Observability best practices to find and fix gaps in our observability and our overall systems. You will also lead practices such as load testing, capacity planning, game days, chaos testing, and incident post-mortems. What You Will Do: Embed within the Engineering pillar to deeply understand the product and implement observability across all key flows Facilitate and build load testing cases, ensuring we understand the limits and scaling factors of our services and systems Contribute to observability and support the design of new services and systems, ensuring highly reliable and scalable concepts are implemented Build and lead practices such as game days, chaos engineering, and failure analysis Build long-term capacity plans, with an eye toward reliability and cost-efficiency Who You Are: 6+ experience writing production-grade software in a modern language, such as Java and Python. Strong knowledge of distributed systems concepts (think CAP theorem), microservices architecture, and distributed tracing . Experience with modern observability systems such as Datadog. Experience with performance debugging tools and patterns. You should be able to read a f
$135K – $207K/yr
Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: As Senior Platform Operations Manager, you will lead the strategy, execution, and optimization of our marketing technology ecosystem. You will architect, implement, and manage the systems and integrations that power our go-to-market (GTM) engine, with a sharp focus on scalability, data integrity, automation, and lead orchestration. This role is pivotal in ensuring that marketing, s
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team focused on transforming how Mastercard's payment systems are built, scaled, and operated. As a Senior Software Engineer, you will lead the design and development of cloud-ready applications, microservices, and distributed systems that support large-scale payment processing platforms while helping advance modernization, automation, and engineering excellence across the organization. In this role, you will contribute to software architecture decisions, drive technical design discussions, and partner with engineers to deliver scalable, resilient, and maintainable software solutions. You'll have the opportunity to solve complex technical challenges, mentor other engineers, and influence how software is designed, developed, tested, and supported across critical technology platforms. What You Will Do •Design software solutions and contribute to software architecture decisions that support scalability, maintainability, and operational excellence. •Translate complex product requirements into technical designs and implementation plans. •Lead development of modular, extensible, high-performance applications. •Design and implement comprehensive unit, functional, and integration testing strategies. •Analyze, optimize, and improve application performance, scal
Other cities to consider
More places hiring for this role
Get new senior system architect jobs in United States by email
Daily job updates · Unsubscribe anytime