Alfred-40B, an OSS RLHF version of Falcon40B
lighton.ai
lighton.ai
We ditched most of our focus on Falcon 40B after Llama 2 70B came out, both the tokens per sec and quality of results are not even close.
Falcon-40B is 63.4 or 61.5 on the non instruction tuned version.
https://github.com/cmp-nct/ggllm.cpp
But its going to be slow without even a small Nvidia GPU (a 2060?). CPUs are really slow at prompt ingestion, and that can't be hidden with streaming.
It uses the ggml library, just like llama.cpp does, and is indeed a fork of llama.cpp's implementation of ggml.
The article was mildly annoying because of many underlines sentences and words that looked like links.