Post

GPU I/O(1) β€” RDMA/GDS

GPU I/O(1) β€” RDMA/GDS

πŸ” Understanding GPU-Centric Data Movement

β€œThe fastest GPU is useless if data cannot reach it efficiently.”
This post traces the evolution of GPU data movement layers and explains how modern architectures minimize CPU involvement using RDMA, GPUDirect, and GDS.


β‘  GPU Data Movement Layer History

Early GPU workloads were dominated by CPU-centric data paths:

1
Storage β†’ CPU Memory β†’ GPU Memory

This model introduced:

  • Redundant memory copies
  • CPU interrupts
  • Cache pollution
  • Latency amplification

To remove these bottlenecks, data movement evolved across three major axes:

GPU Data Movement History Figure 1. GPU Data Movement History.

  1. RDMA (1999+)
    • NIC ↔ Host Memory
    • Zero-copy across hosts
  2. GPUDirect RDMA (2012+)
    • NIC ↔ GPU Memory
    • GPU-visible memory mapping
  3. GDS (2020+)
    • NVMe ↔ GPU Memory
    • Storage bypasses CPU entirely

Each layer progressively removes the CPU from the data path.


β‘‘ GPU Data Movement Architecture

Modern GPU data movement is built on DMA-capable endpoints:

  • GPU
  • NIC
  • NVMe SSD

All of them can act as bus masters on PCIe.

Core Principle

If two devices can perform DMA and share addressability, the CPU does not need to touch the data.

Unified View (Local + Remote)

1
2
GPU Memory ↔ NIC ↔ Network ↔ NIC ↔ GPU Memory
GPU Memory ↔ NVMe (PCIe DMA)

The CPU’s role is reduced to:

  • Queue setup
  • Descriptor submission
  • Control-plane orchestration

No payload data passes through CPU caches.


β‘’ GDS Data Movement (Local)

GPUDirect Storage (GDS) enables direct data transfer:

GPU Local Data Movement Figure 2. GPU Local Data Movement.

1
Local NVMe SSD β†’ GPU Memory

Without GDS

1
NVMe β†’ System Memory β†’ GPU Memory
  • 2 DMA hops
  • CPU page cache involvement
  • Higher latency

With GDS

1
NVMe ──DMA──▢ GPU Memory

Characteristics:

  • Single DMA operation
  • No CPU copy
  • No cache pollution
  • Deterministic latency

GDS treats GPU memory as a first-class I/O target.


β‘£ Remote GDS Data Movement

Remote GDS extends the same principle across hosts.

GPU Remote Data Movement Figure 3. GPU Remote Data Movement.

Data Path

1
2
3
4
5
Remote NVMe
 β†’ Remote NIC
 β†’ RDMA Fabric
 β†’ Local NIC
 β†’ GPU Memory

Key technologies involved:

  • NVMe-oF (RDMA)
  • GPUDirect RDMA
  • GDS

β‘€ TL;DR

  • RDMA removed CPU from network data paths
  • GPUDirect RDMA extended RDMA to GPU memory
  • GDS removed CPU from storage I/O
  • Remote GDS combines both:
    • NVMe-oF + RDMA + GDS
    • GPU ↔ Storage, end-to-end, zero-copy
This post is licensed under CC BY 4.0 by the author.