Jobs in United States

Senior Ai Compute Engineer in United States

1,941 active opportunities · Updated October 2026

Explore current senior ai compute engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule

G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a highly experienced RTL engineer to own critical on- and off-chip interconnect components for our custom AI accelerator platform. You will drive the microarchitecture and RTL implementation of scalable on-chip communication fabrics connecting high-bandwidth compute, memory, and I/O subsystems as well as purpose-built off-chip interfaces and protocols needed to enable custom computing at scale. This is a senior, hands-on engineering role with broad technical ownership. You will drive design from requirements through the full silicon lifecycle, from architecture definition and performance analysis through RTL implementation, verification closure, physical design convergence, bring-up, and production readiness. You will plan and oversee the work of junior engineers and help drive and develop productive engineering relationships with external partners and help manage partner execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the microarchitecture, RTL design, and delivery of major SoC interconnect components, including network-on-chip fabrics, switches, routers, bridges, protocol adapters, arbiters, and traffic-management logic as well as off-chip protocol bridges and interfaces. Drive third party engagements to develop novel networking and interface protocols and silicon IP while ensuring high quality and de

AWSRestAIRust
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

1671 About the Role We are seeking a Senior Signal Integrity Engineer to develop, validate, and optimize high-speed signaling solutions across blade- and rack-level architectures for advanced compute platforms. This role sits at the intersection of silicon, package, interconnect, board, and system design , with a strong emphasis on hands-on measurement, simulation correlation, and cross-functional technical communication. Key Responsibilities End-to-end signal integrity analysis for blade- and rack-level system architectures. Analyze and optimize high-speed and low-speed I/O interfaces, including PCIe Gen4/5/6, Ethernet, DDR, SerDes, SPI , I2C, etc . and related interconnects. Perform time-domain and frequency-domain simulations using tools such as Ansys HFSS, Keysight ADS, Cadence Sigrity, CST, SPICE , or similar. Support hands-on lab validation using VNA, TDR, BERT, and high-speed oscilloscopes . Correlate simulation results with lab measurements to identify margin gaps, debug issues, and improve design methodology. Collaborate with silicon, package, board, connector, cable, and system design teams to optimize I/O channel performance. Work with interconnect vendors and ODMs to guide board layout, stack-ups, routing rules, and system design decisions. Review schematics, layouts, simulation results, and validation data for blade, backplane, and rack-level hardware. Prepare and communicate clear validation reports, measurement summaries, debug findings, and technical recommendations to internal teams, vendors, and senior technical stakeholders. Required Qualifications Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, or a related field . Strong experience in signal integrity for high-speed digital systems. Hands-on measurement expertise using VNAs, TDRs, BERTs, and high-speed oscilloscopes . Experience measuring and analyzing S-parameters, impedance profiles, eye diagrams, jitter, timing margins, insertion loss, return loss, and cro

PythonGitAI
S
📍 Bellevue, Washington, United States· Full-time· Remote
✓ High-confidence listingCompany trend -91.7%
Quick readStrong listing-quality and freshness signals

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

KubernetesAIGoRust
R
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At Render, we’re building the modern cloud platform for developers creating AI-native, full-stack, multi-service applications. Our mission is to eliminate the tradeoff between the power of hyperscalers and the simplicity of developer-friendly platforms—so teams can ship fast, scale reliably, and focus on their product, not infrastructure. Unlike complex hyperscalers or ephemeral edge/serverless solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability, allowing teams to launch, scale, and manage real-world applications without writing infrastructure code or managing servers. Whether you're building LLM-powered applications, scalable SaaS products, or async processing pipelines, Render empowers teams to move fast and scale confidently from MVP to millions of users. Our platform is trusted by over 6.5 million developers worldwide and continues to grow rapidly. In February 2026, we raised an additional $100M in Series C financing, bringing our total funding to $257M, to accelerate our vision of making cloud infrastructure both powerful and intuitive—designed for the speed of modern AI development. We’re a diverse and talented team that values craft, velocity, and user experience. If you’re excited to help shape the future of the intelligent cloud and empower developers everywhere, we’d love to hear from you. Applying to Render We're seeking candidates who possess high integrity, humility, and an insatiable drive to learn. Through reasoned discussions and continuous feedback, we strive to improve both individually and collectively. We foster an environment of mutual trust and respect, empowering effective debate to achieve the best outcomes for our customers and team. We especially encourage members of underrepresented groups in the tech community to apply and understand that not all successful candidates will meet each requirement listed. Our interview process is unique to each role, an

RestAIGoRust
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -91.7%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Our team builds and operates Snowflake's Streaming Platform — the core infrastructure responsible for bringing data into Snowflake continuously and at scale. This includes Snowpipe Streaming and Datastream, services used by large set of enterprise customers to move mission-critical data in real time. Data is the fuel powering the new enterprise AI world, and we are the team responsible for getting it there reliably, continuously, and fast. Responsibilities Design, build, and maintain core components of Snowflake's streaming ingestion platform Improve service reliability, scalability, and latency under high-throughput production workloads Debug and resolve incidents in a distributed, multi-tenant cloud service Contribute to the design and evolution of streaming APIs, SDKs, and server-side protocols Write comprehensive tests including unit, integration, and chaos/fault-injection scenarios Collaborate with partner teams (storage, query, compute) on cross-cutting platform concerns Participate in code reviews, on-call rotations, and architecture discussions Required Qualifications 3–5 years of software engineering experience on large-scale distributed systems or cloud services Strong proficiency in Java or C++ Deep understanding of distributed systems concepts: consistency, faul

JavaAIC++Go
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -91.7%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

KubernetesAIGoRust
R
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At Render, we’re building the modern cloud platform for developers creating AI-native, full-stack, multi-service applications. Our mission is to eliminate the tradeoff between the power of hyperscalers and the simplicity of developer-friendly platforms—so teams can ship fast, scale reliably, and focus on their product, not infrastructure. Unlike complex hyperscalers or ephemeral edge/serverless solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability, allowing teams to launch, scale, and manage real-world applications without writing infrastructure code or managing servers. Whether you're building LLM-powered applications, scalable SaaS products, or async processing pipelines, Render empowers teams to move fast and scale confidently from MVP to millions of users. Our platform is trusted by over 6.5 million developers worldwide and continues to grow rapidly. In February 2026, we raised an additional $100M in Series C financing, bringing our total funding to $257M, to accelerate our vision of making cloud infrastructure both powerful and intuitive—designed for the speed of modern AI development. We’re a diverse and talented team that values craft, velocity, and user experience. If you’re excited to help shape the future of the intelligent cloud and empower developers everywhere, we’d love to hear from you. Applying to Render We're seeking candidates who possess high integrity, humility, and an insatiable drive to learn. Through reasoned discussions and continuous feedback, we strive to improve both individually and collectively. We foster an environment of mutual trust and respect, empowering effective debate to achieve the best outcomes for our customers and team. We especially encourage members of underrepresented groups in the tech community to apply and understand that not all successful candidates will meet each requirement listed. Our interview process is unique to each role, an

RestAIGoRust
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li

PythonDockerKubernetesArtificial Intelligence
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -84.3%

From $208.6K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Intro : Pinterest’s Service Configuration & Coordination team builds the platforms, APIs, and tooling that make dynamic configuration and service coordination safe, scalable, and reliable for every service across regions and clouds, serving as the central hub for cross-functional alignment and system orchestration. As a Senior Staff Software Engineer (IC17) , you’ll own the end-to-end technical strategy and long-term vision for our configuration and coordination infrastructure. You’ll define the future of these mission-critical systems, partnering closely with Traffic, Compute, Observability, and Cloud Architecture teams to deliver robust, high-performance solutions at scale. What you’ll do: Orchestrate the long-term technical vision for configuration ecosystems to ensure feature flags, ML settings, and experiments utilize unified pave

AWSCI/CDRestAI
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -91.7%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Senior Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A SENIOR SOFTWARE ENGINEER, YOU WILL: ● Design and develop features, understand customer requirements and meet business goals. ● Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. ● Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. ● Build highly reliable and fault-tolerant software to meet the needs of the largest customers. ● Ensure operational readiness and ma

JavaVueAWSAzure
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -91.7%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Senior Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A SENIOR SOFTWARE ENGINEER, YOU WILL: ● Design and develop features, understand customer requirements and meet business goals. ● Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. ● Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. ● Build highly reliable and fault-tolerant software to meet the needs of the largest customers. ● Ensure operational readiness and ma

JavaVueAWSAzure
🔔

Get new senior ai compute engineer jobs in United States by email

Daily job updates · Unsubscribe anytime