Micron Addresses Memory Bottlenecks at Hot Chips; Samsung PIM Speeds Up LLM Inference by 3x
Key point
At Hot Chips 2026, Micron analyzed the gap between compute and memory bandwidth and revealed that Samsung's LPDDR5X PIM technology improved Llama 3.1 8B inference speed by 3.01x.
Details
At Hot Chips 2026, Micron visualized the gap between compute performance and memory bandwidth through Raghu Sreeramaneni's memory tutorial. Normalized TFLOPS from TPU v3 to R200 increased by approximately 3x every two years, while bandwidth from HBM2e to HBM4 increased by less than 2x, significantly widening the disparity between the two metrics.
To address this 'Memory Wall' issue, three approaches currently being implemented by the industry were introduced:
- Memory placement next to compute: Reducing physical distance
- Proximity placement via short links: Minimizing connection latency
- Processing-in-Memory (PIM): Performing calculations directly within memory without data movement
Notably, regarding PIM technology, Samsung applied it to LPDDR5X and presented measurement results showing a 3.01x improvement in tokens generated per second during Llama 3.1 8B model inference.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.