164 karma · joined May 27, 2023
It's getting tougher to use older, cheaper GPUs (Pascal/Maxwell) with modern quantization schemes so anything you can do to keep kernels compatible with SM52 and SM61 would be greatly appreciated.
You can take your data to another server only if the original server is up. When the admin forgets to pay the bill and ISP nukes the box and the backups admin said they had don't exist and the coadmin turns out to have no access at all, you lose everything. Ask me how I know.
Self hosting in theory should fix this but it doesn't: it's a resource intensive afterthought because the developers only really care about their Big Instance.
If your usecase fits inside that 32GB (no 70B models, sadly) the price to performance of a GGUF Q4KM is really attractive on this setup.
I've had good luck with OneProvider.
If you have Nvidia hardware the ctranslate2 based faster-whisper is very very fast: https://github.com/guillaumekln/faster-whisper
The dropdown at the top selects which comparison: Falcon compares GGML, Vicuna compares bits and bytes. I have some more comparisons planned, feel free to open an issue if you'd like to see something specific: https://github.com/the-crypt-keeper/can-ai-code
What causes such loops? Just a challenge over and over.