2 Matching Annotations
  1. Last 7 days
    1. Prefilling is the first step, which processes the input prompt to generate the initial KV (key-value) cache values in addition to the first generated token. This phase is parallelized but compute-bound (Delavande et al. 2026), which means it is limited by pure mathematical calculation speed on the accelerator and usually leads to saturation in terms of compute and power
    2. This process is memory-bound, which means it is limited by how fast the hardware can stream model weights and cache data from memory. By default (i.e., with no optimization), this usually leads to low utilization, with a GPU generally running at 30-60% of its total capacity in terms of power draw and memory utilization

      I have not realised that the decode phase was so much less energetically expensive