Super curious how you did this! Doesn’t 30B model require a hefty computer to run locally (assuming you’re tuning a non-quantized version)
What I'm wondering is how they fed the documents, as all those LLMs have limitations on the input sizes.
> What I'm wondering is how they fed the documents, as all those LLMs have limitations on the input sizes.
It's like file hashing at scale, you don't have to read the whole stream for every file, just the first 1024/2048 bytes (or first few paragraphs).
(This works for classification and sorting, less so for summarization.)
That's for the 7B model. The 30B model needs 24GB quantized (or 64GB for the unquantized model).