Really interesting is that current LLMs don’t explicitly use chunking for storage; they rely on distributed representations across parameters. However, their self-attention mechanisms and sequence processing mimic chunking during runtime, creating dynamic "chunks" of context.
I’m at my limit, wondering if future models incorporating explicit chunking for better memory, scalability, and efficiency could truly take them to the next level.