CFM has updated the domain of its english website from en.chinaflashmarket.com to www.memorymarket.com, Please be informed.

Micron Reportedly Explores Near-GPU NAND Flash Development to Boost AI Workload Performance

By: M 1 day ago

According to reports, Micron is exploring "near-GPU NAND," a new type of high-endurance NAND Flash memory placed closer to GPUs to act as a fast, cost-effective buffer between GPU memory and traditional storage. This innovative approach, often referred to as "Near-GPU NAND," aims to provide a compromise between memory density and endurance, creating an optimized new memory tier for modern computing demands.

The concept behind Near-GPU NAND is to offer a storage solution that brings NAND flash closer to the GPU and can handle workloads that do not require the low latency or high bandwidth of HBM or conventional DRAM modules. This approach allows GPUs to access a pool of NAND flash with lower density compared to traditional TLC or QLC NAND, but with a focus on improving I/O speed, bandwidth, and read performance.

For AI workloads, especially those involving large-scale LLMs (Large Language Models), Near-GPU NAND is expected to alleviate memory bottlenecks. This solution enables GPUs to directly access a potential storage pool of hundreds of gigabytes, similar to the existing capabilities of DRAM such as HBM or GDDR7/LPDDR6. This tier would sit between GPU memory and regular system storage, acting as a very fast buffer, which would be ideal for the inference of massive LLMs. At that point, LLMs wouldn't be memory-bound and could run on systems with fewer GPUs, but with more storage and memory tiers to compensate for the required inference size.

Furthermore, by leveraging expanded storage and memory tiers, systems could run complex models with fewer GPUs, reducing the demand for expensive HBM or DRAM. NAND flash would be more cost-effective than HBM or regular DRAM, and this approach not only optimizes performance but also provides a more cost-efficient solution that can be tailored to specific system requirements.