Inside PyTorch(4) — How to Analyze GPU Runtime Performance
🔍 Once a Model Is Compiled, How Do We Know Whether the GPU Is Actually Running Efficiently?
After compilation, performance analysis should focus on the steady-state runtime: establish a reproducible baseline, identify where GPU time is lost, locate expensive kernels, verify whether efficient implementations were selected, and finally trace inefficient kernels back to compiler decisions when necessary.
① Establish a Reproducible Performance Baseline
Before profiling, define the workload conditions that will remain fixed during optimization: representative input shapes, batch size, precision, warm-up iterations, concurrency, CUDA Graph usage, and measurement duration. Separate compilation or engine-build time from steady-state execution time, and record baseline latency and throughput before making any changes.
② Check Basic GPU Telemetry
Use nvidia-smi dmon or NVIDIA DCGM to check GPU/SM utilization, HBM activity, clocks, power, temperature, PCIe, and NVLink activity. This provides the first indication of whether the GPU is underutilized or throttled.
③ Profile PyTorch Execution
Use torch.profiler and NVTX to identify expensive forward, backward, optimizer, dataloader, and operator regions before moving to lower-level GPU analysis.
④ Inspect the CPU-GPU Timeline
Use NVIDIA Nsight Systems (nsys) to find GPU idle gaps, delayed kernel launches, synchronization, memory copies, and CPU-side stalls. Large gaps between GPU kernels usually indicate that the GPU is waiting for work.
⑤ Inspect Multi-GPU Communication
For distributed workloads, use Nsight Systems with NCCL tracing and DCGM to inspect AllReduce, AllGather, ReduceScatter, rank imbalance, and compute-communication overlap.
⑥ Inspect GPU Memory Behavior
Use Nsight Systems GPU Memory tracing and torch.profiler to inspect H2D/D2H transfers, allocation/deallocation behavior, peak memory usage, and unnecessary memory movement.
⑦ Investigate Unexplained GPU Idle Time
If GPU gaps cannot be explained by CPU, communication, or memory activity, use tlparse and TORCH_LOGS to inspect graph breaks, eager fallbacks, recompilations, fusion, and scheduling decisions.
⑧ Identify Hotspot GPU Kernels
Use Nsight Systems or torch.profiler to rank kernels by total GPU execution time. Focus detailed analysis on the small number of kernels responsible for most of the runtime rather than profiling every kernel.
⑨ Inspect Algorithm and Kernel Variant Selection
A GPU operation often has multiple valid implementations with different tiling, scheduling, memory usage, workspace requirements, and hardware efficiency. Before manually optimizing a hotspot kernel, verify whether the runtime, compiler, or library selected an efficient implementation for the actual workload.
Some systems use heuristics, while others benchmark multiple candidates and select the fastest one. TensorRT profiles alternative tactics during engine building, TorchInductor max-autotune benchmarks Triton or template implementations, and GPU libraries such as cuDNN may select among multiple algorithms. Treat these mechanisms as a general implementation-search layer between graph optimization and individual kernel analysis.
For dynamic workloads, perform tuning with representative shapes rather than assuming that a configuration selected for one shape remains optimal for every shape. Reuse timing or tuning caches only when their hardware and software assumptions remain valid.
🔗 NVIDIA TensorRT — Optimizing Performance 🔗 PyTorch — torch.compile max-autotune 🔗 TorchInductor GPU Profiling
⑩ Analyze Individual Kernel Efficiency
Use NVIDIA Nsight Compute (ncu) on hotspot kernels. Check SpeedOfLight/Roofline, compute and memory throughput, Tensor Core utilization, occupancy, registers, shared memory, active/eligible warps, cache/HBM behavior, and warp stall reasons.
⑪ Trace Kernel Problems Back to the Compiler
If a kernel shows poor hardware efficiency or no efficient implementation is selected, return to tlparse and Inductor logs such as fusion, schedule, ir_pre_fusion, ir_post_fusion, and kernel_code. Check whether excessive fusion, insufficient fusion, scheduling, generated Triton code, or the available implementation search space explains the observed GPU behavior.
⑫ Validate Every Optimization
After changing batch size, fusion behavior, precision, CUDA Graphs, compiler options, algorithm selection, or autotuning configuration, repeat the same benchmark and profiling workflow. Compare throughput, latency, GPU idle time, hotspot kernel time, and hardware efficiency while also verifying numerical correctness. Measure build or compilation overhead separately from steady-state runtime improvements.
Key Takeaway
GPU performance optimization is not only about finding slow kernels. It should proceed from reproducible system-level measurement toward hotspot identification, implementation selection, individual kernel analysis, and finally compiler decisions when necessary.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
Reproducible Baseline
↓
DCGM / nvidia-smi
↓
torch.profiler + NVTX
↓
Nsight Systems
↓
┌────┴────┐
▼ ▼
GPU Idle GPU Busy but Slow
│ │
▼ ▼
tlparse Hotspot Kernels
↓
Algorithm / Tactic
/ Autotune Search
↓
Nsight Compute
↓
Inductor / Triton
Analysis
↓
A/B Validation
TL;DR
Nsight Systems: Where is runtime being lost?Nsight Compute: Why is a specific GPU kernel inefficient?tlparse/Inductor logs: Why did the compiler generate that execution structure?DCGM/torch.profiler: Isolate the problem before going deeper.