Post

CUDA(3) — AABS

CUDA(3) — AABS

🔍Understanding Allocated Active Blocks per SM (AABS)

① What Is AABS?

AABS (Allocated Active Blocks per SM) refers to the number of thread blocks that are simultaneously allocated and actively running within a single Streaming Multiprocessor (SM).
It represents how the GPU hardware distributes parallel work units (blocks) across SMs.

Formula:
AABS = min(BlockLimit, RegisterLimit, SharedMemoryLimit, WarpLimit)

That means the number of active blocks per SM is determined by the most restrictive resource among the four limits.


② Resource Constraints

The ncu Command Result Example Figure 1. Example result of the ncu (Nsight Compute) command on NVIDIA RTX A6000 (GA102, Ampere Architecture).

Limiting FactorDescriptionExample (A100)
Block limitArchitectural maximum number of blocks per SM32 blocks
Register limitLimited register file per SM, affected by registers per thread65,536 registers total
Shared memory limitSM shared memory divided among active blocks164 KB total
Warp limitDepends on block size (warps per block × blocks per SM ≤ 64)64 warps

The smallest of these values defines how many blocks can coexist on an SM.


③ Relationship to Occupancy

Occupancy is the ratio of active warps to maximum warps per SM:

Occupancy = Active Warps / Max Warps per SM

Since each block contains several warps, the number of active blocks directly determines active warps — hence occupancy.

1
2
3
4
5
6
7
Registers / Shared Memory / Warp Limit / Block Limit
                ↓
     Allocated Active Blocks per SM
                ↓
         Active Warps per SM
                ↓
          Achieved Occupancy

In short: resource limits → blocks → warps → occupancy.


④ Example Calculation

For NVIDIA A100:

  • Total registers: 65,536 per SM
  • Shared memory: 164 KB per SM
  • Max warps per SM: 64
  • Block size: 256 threads (8 warps)
  • Each thread uses 48 registers
  • Each block uses 24 KB shared memory
FactorCalculationLimit
Registers65,536 / (256×48) ≈ 5 blocks5
Shared Memory164 KB / 24 KB = 6 blocks6
Warp64 / 8 = 8 blocks8
Block Limit32 blocks32
Resultmin(5,6,8,32) = 5 blocks/SM✅

Thus:
Active Warps = 5 × 8 = 40,
Occupancy = 40 / 64 = 62.5%


⑤ Why It Matters

Higher AABS values mean more concurrent blocks → more active warps → higher potential occupancy.
However, just like occupancy, more is not always better — excessive register or memory pressure may cause performance drops.

⚠️ Optimizing AABS involves balancing block size, register usage, and shared memory to maximize useful concurrency.


Summary

ConceptDescription
AABSNumber of active blocks allocated per SM
Determined byMinimum of {Block, Register, Shared Memory, Warp limits}
Related toOccupancy (active warps / max warps)
GoalMaximize concurrency without resource exhaustion
ToolNsight Compute → sm__active_blocks_per_multiprocessor

References

  • NVIDIA CUDA Programming Guide, v12.4
  • Nsight Compute Metrics: sm__active_blocks_per_multiprocessor, sm__warps_active.avg.pct_of_peak_sustained_active
  • NVIDIA A100 Whitepaper — Architecture and Resource Partitioning
This post is licensed under CC BY 4.0 by the author.