1,008 karma · joined April 3, 2019
On the other hand, the main points of Zamba/Mamba are low latency, generation speed, and efficient memory usage. If this is true, LLMs could be much easier for everyone to use. All we need to do is wait for someone with a good training dataset to train a SOTA Mamba.
How can you retrieve the latent representation of the candidate LLMs? Some models do not have open weights (such as GPT-4), which means AFAIK it is impossible to directly access the hidden latent space through their API.
Am I missing something?
Does this mean LLMs can generate text with empty context? How can LLMs choose the first token without any previous tokens? My understanding is that to compute logits for the next token, LLMs require input from all previous tokens. Am I correct?