If you just want to play around with an LLM though, absolutely.
If you just want to play around with an LLM though, absolutely.
Between that and dirt cheap storage prices, it is possible to have a local, offline copy of more human knowledge than one can sensibly consume in a lifetime. Hell, it's possible to have it all on one's smartphone (just get one with an SD card slot and shove a 1+ Tb one in there).
> on an old raspberry pi
I bet the LLM responses will be great... You're better off just opening up a raw text dump of Wikipedia markup files in vim.
First there is a leaderboard for embeddings. [1]
Even then, it depends how you use them. Some embeddings pack the highest signal in the beginning so you can truncate the vector, while most can not. You might want that truncated version for a fast dirty index. Same with using multiple models of differing vector sizes for the same content.
Do you preprocess your text? There will be a model there. Likely the same model you would use to process the query.
There is a model for asking questions from context. Sometimes that is a different model. [2]
Are LLMs unable to distinguish between real life and fantasy? What prompts have you thrown at them to make this determination? Sending a small fairy tale and asking the LLM if it thinks it's a real story or fake one?
"Shit in, shit out" as the saying goes, but applied to conversations with LLMs where the prompts often aren't prescriptive enough.