293 karma · joined August 26, 2021
lood_in_4bit=True will let you run Llama2-7B variants at 6.3GB VRAM.
Thanks to this post, I kind of have another idea for a 'corrections to transcript' feature that Llama2-7B even on CPU can help with.
As for the former classifier you can try doing zero-shot classification between n number of categories + others. Models like Flan-T5/T5/Flan-UL2/DistillBART(also ~7B-40B param LLMs can also do this but would be overkill).
I'm trying to evaluate best serverless solutions for inference without compromising on client usage & reducing idle time on GPU boxes. So far its down to base10, HF, Banana, I'll end up pooling them all & then sending requests between them. For dedicated training boxes Lambda, Modal, Oblivus, Runpod are the contenders.
My downloads folder is effectively categorized as
AW - artwork, could be pics, videos that I create or download, further into monthly phone backup, jellyfin etc
SW - software, contains anything I've downloaded, specific folders for each caterory like 3D printing, setup files, repo code zips/model weights etc
PDFs - as the name implies PDFs only, mostly arxiv papers grouped by month-year folders & backed off monthly to phone & NAS
-----
Has worked like a charm for me for a while & very easy to take backups of my machine since everything essential is present within the Download sub-folders