First domestic model (Union?)
28 karma · joined January 16, 2024
Instead for decode, you need to sequentially generate each token.
Decode is the next major step where you start generating output tokens one at a time.
Both run on GPUs but have slightly different workloads
1. Prefill has very little I/o from VRAM to HBM and more compute 2. Decode is light on compute but have to I/o the keys and values computed in the prefill stage for every output token