GPU Memory(1) — Architecture
🔍Understanding the GPU Memory Network
In the previous post,
we traced how commands travel — from the CUDA driver to SMs.
Now, to understand how operands (data) move,
we must first understand the GPU memory network that physically delivers that data.
Because no instruction runs without its operands,
understanding the data path is just as essential as the command path.
① Big Picture — The Memory Highway
Figure 1. NVIDIA A100 GPU Architecture (MIG Example).
Figure 2. Memory hierarchy in GPUs.
From compute cores down to physical DRAM:
1
SMs → L2 Cache → Memory Controllers → HBM Stacks
Every byte fetched by a CUDA kernel flows through this route.
Each layer adds coordination, bandwidth, and parallelism.
② Streaming Multiprocessor (SM) — Where Computation Starts
1
SM = { CUDA Cores + Tensor Cores + Warp Scheduler + L1 Cache + Registers }
- The SM (Streaming Multiprocessor) executes kernel instructions.
- Each SM runs thousands of threads in groups of 32 (warps).
- When a warp stalls on memory, another warp is scheduled — maximizing utilization.
- The L1 cache and shared memory sit closest to execution units.
- Operands first come from registers, then from L1, then from global memory (via L2).
③ GMMU & L2 Cache — Translating and Buffering
1
SM → GMMU → L2 Cache → Memory Controller
- The GMMU (GPU Memory Management Unit) converts virtual addresses (GPU VA) into physical addresses (PA).
- After translation, the data request hits L2 Cache, which is shared by all SMs.
- The L2 acts as a crossroad — merging and reordering requests before they go to memory controllers.
- On a GPU like NVIDIA A100, the L2 is divided into slices connected to each controller via a high-speed on-die fabric.
④ Memory Controller — The Traffic Dispatcher
1
Controller count = Channel count
- Each Memory Controller corresponds to one HBM channel.
- It translates memory requests into DRAM-level commands (
ACT,READ,WRITE,PRE). - It handles timing for bank groups and banks,
ensuring proper precharge and activation cycles. - Multiple controllers operate in parallel — achieving aggregate terabyte-level bandwidth.
⑤ High-Bandwidth Memory (HBM) — The Physical Store
Figure 3. HBM Stacked DRAM Architecture.
1
2
3
4
5
HBM Stack
├─ Channel 0–7
│ ├─ BankGroup 0–7
│ │ └─ Banks (DRAM cells)
└─ TSV links (vertical connections)
- Each HBM stack has 8 independent channels, each 128-bit wide.
- 8 × 128 bits = 1024-bit aggregate bus per stack.
- Each channel has its own command/address bus and I/O interface.
- Inside each channel are bank groups, each containing multiple banks (storage arrays).
- Data moves vertically between DRAM layers through TSVs (Through-Silicon Vias).
⑥ Data Path — From Compute to Silicon
1
2
3
4
5
6
7
8
9
SMs (Compute)
↓
GMMU (Address Translation)
↓
L2 Cache (Shared Buffer)
↓
Memory Controller (Command Control)
↓
HBM Channel → BankGroup → Bank (Physical Storage)
- SMs generate load/store requests.
- GMMU maps virtual to physical memory.
- L2 caches and merges requests from all SMs.
- Memory Controllers issue low-level DRAM commands.
- HBM stores the actual operands that kernels operate on.
Each layer contributes to latency hiding and bandwidth scaling —
a vertical network that bridges software addresses and physical electrons.
⑦ Layer Summary
| Layer | Role | Parallelism Unit |
|---|---|---|
| SM | Compute instructions | Warp |
| GMMU | Address translation | Page |
| L2 Cache | Shared buffer | Slice |
| Memory Controller | Command scheduling | Channel |
| HBM | Physical storage | BankGroup / Bank |
⑧ Why This Matters
In the previous article, commands told the GPU what to do.
In this one, we explored where the data lives and how it moves when those commands execute.
Every
ld.globalorst.globalinstruction inside a kernel ultimately traverses
the GMMU → L2 → Controller → HBM hierarchy you just saw.
Understanding this path lays the groundwork for the next post,
where we’ll connect both sides — command flow and data flow —
to see how computation and memory converge on the GPU. ⚡️