Post

GPU Memory(2) β€” Data Paths

GPU Memory(2) β€” Data Paths

πŸ”GPU Memory Data Paths β€” How Bytes Actually Move

In the previous post,
we explored the structure of the GPU’s memory system β€” how SMs, caches, controllers, and HBM stacks connect.

PCIe Switchs Fabric Figure 1. PCIe Switchs Fabric.

Now it’s time to follow the actual movement of data through that structure.
This post focuses on how bytes travel across the PCIe fabric, BAR mappings, and DMA engines
β€” from CPU or NIC all the way into GPU VRAM and HBM.

Understanding these paths completes the picture:
if the last article explained where data can live,
this one shows how it actually gets there.


β‘  PCIe Hierarchy Overview

1
2
3
4
5
Host CPU
 └─ Root Complex (RC)
      └─ PCIe Switch
            β”œβ”€ GPU (Endpoint)
            └─ NIC (Endpoint)
  • Root Complex (RC) connects CPU/system memory to the PCIe fabric.
  • PCIe Switch routes transactions between devices (GPUs, NICs, NVMe).
  • Each endpoint has its own BARs (Base Address Registers), exposing parts of its memory to the PCIe bus.

β‘‘ Base Address Register (BAR) β€” The Window to Device Memory

BAR (Base Address Register) defines a memory window that maps a device’s internal memory into the PCIe address space.

PCI BAR memory addresses Figure 2. Device memory regions (BARs) are mapped into the same address space as system memory.

Example for a GPU:

1
2
3
BAR0: MMIO control registers
BAR1: VRAM window (device memory space)
BAR2+: extended or configuration ranges

BAR = a PCIe-visible window into GPU memory, not the memory itself.

When the CPU or another device writes to a BAR address,
it’s actually sending packets over PCIe that get translated into GPU memory accesses.


β‘’ PCIe Switch Fabric β€” The Router of Data

The PCIe Switch routes all transactions based on their target address range.
Each downstream port corresponds to a device’s BAR region.

1
2
3
4
5
PCIe Switch
 β”œβ”€ Upstream Port β†’ Root Complex
 β”œβ”€ Downstream Port 1 β†’ GPU
 β”œβ”€ Downstream Port 2 β†’ NIC
 └─ Downstream Port 3 β†’ NVMe

When the CPU accesses a GPU BAR address,
the switch simply forwards that PCIe Transaction Layer Packet (TLP) to the correct device.

The switch doesn’t know what the data means β€” it only routes packets by address.


β‘£ Local Node Transfers β€” CPU ↔ GPU

Intra-node data movement uses BARs within the same PCIe fabric.

1
CPU β†’ PCIe Root Complex β†’ PCIe Switch β†’ GPU (BAR1 β†’ VRAM)
  1. The CPU issues a DMA or MMIO write to a GPU BAR address.
  2. The PCIe controller wraps it into TLPs (Transaction Layer Packets).
  3. The switch forwards them to the GPU port.
  4. The GPU decodes the BAR address, resolves it to a VRAM offset,
    and its DMA engine performs the actual write into VRAM.

The CPU never moves the bytes itself β€” it just tells the DMA engine where to move them.

This mechanism is what enables GPUDirect Storage (GDS) β€”
NVMe devices can DMA directly into GPU VRAM via BAR1, bypassing host memory entirely.


β‘€ Peer-to-Peer Transfers β€” GPU ↔ GPU (Same Node)

Modern GPUs can also communicate directly over PCIe without CPU involvement.

1
GPU A (BAR) ↔ PCIe Switch ↔ GPU B (BAR)
  • Each GPU exposes its memory window through BAR1.
  • The initiating GPU uses its DMA engine to write directly into the peer’s BAR address.
  • The switch routes the transaction between GPUs.
  • CUDA’s peer-to-peer (P2P) and NCCL backends rely on this mechanism.

BAR-to-BAR transfers are the foundation of GPUDirect P2P communication.


β‘₯ RDMA Transfers β€” GPU ↔ GPU (Across Nodes)

For inter-node transfers, the NIC becomes part of the data path.
The NIC’s DMA engine accesses GPU memory via BAR, while another NIC mirrors the process remotely.

1
2
3
4
5
GPU A (VRAM)
   ↑
PCIe ←→ NIC A ──── RDMA Fabric ──── NIC B ←→ PCIe
                                               ↓
                                         GPU B (VRAM)

Steps:

  1. The CPU driver (e.g., mlx5_core) registers GPU memory and exposes it as a Remote BAR (RKey).
  2. The NIC’s DMA engine performs PCIe reads/writes directly into that BAR region.
  3. The remote NIC does the same, completing GPU↔GPU transfers with no CPU copy.

RDMA turns the BAR window into a globally addressable GPU memory region β€” enabling true zero-copy between nodes.


⑦ The Full Data Path

1
2
3
4
5
6
7
8
9
Host Memory / Storage
   ↓
PCIe Root Complex
   ↓
PCIe Switch Fabric
   ↓
GPU BAR (PCIe-visible address)
   ↓
GPU DMA Engine β†’ Memory Controller β†’ HBM (Physical Data Store)
  • CPU or NIC initiates a PCIe transaction to the GPU BAR address.
  • PCIe Switch routes the TLPs to the GPU.
  • GPU’s DMA Engine writes data into VRAM.
  • The Memory Controller inside the GPU moves it into HBM through internal DRAM channels.

Thus, bytes flow from system or remote memory β†’ PCIe fabric β†’ GPU BAR β†’ HBM stack.


β‘§ Why It Matters

In the previous post,
we saw how commands travel to tell the GPU what to do.
In this post, we examined how data actually travels to where it’s needed.

Together with GPU Memory Architecture,
these layers complete the story of how commands and data meet inside the GPU.

Command Path = PFIFO β†’ PBDMA β†’ SMs
Data Path = PCIe BAR β†’ DMA β†’ Memory Controller β†’ HBM


Summary

Path TypeInitiatorMechanismExample
CPU ↔ GPUCPU DMAPCIe BARGPUDirect Storage
GPU ↔ GPU (local)GPU DMABAR-to-BARNCCL / CUDA P2P
GPU ↔ GPU (remote)NIC DMARDMA (Remote BAR)GPUDirect RDMA
This post is licensed under CC BY 4.0 by the author.