1. Binary build https://github.com/jaykrell/llama.cpp/releases/tag/1
2. Quantized model (7B/13B/30B) https://mega.nz/folder/UjAUES6Z#bGhKkyiZX3eRrn9HcxVVfA
3. main.exe -m ggml-model-q4_0.bin -t 8 -n 128
main.exe -m ggml-model-q4_0.bin -t 8 -n 128 -p "The Drake equation is nonsense because"
The Drake equation is nonsense because it takes parameters that can only be known AFTER the conclusion is reached. It would be like saying "I'm going to prove a theorem by starting from the conclusion, then making up the proof. The Drake equation uses the existence of extraterrestrial intelligence as the conclusion and then making up the parameters. It is nonsense.
But, quantize.exe doesn't seem to work - any valid command (such as below) pauses for a split second, then returns with no output?
$ quantize.exe ggml-model-f16.bin ggml-model-q4_0.bin 2
I wasn't sure where to upload them, and that link is only good for 50 downloads. Can put them somewhere else if you know a better location that doesn't require signup.
llama.exe is basically main.exe?
I actually learned how to compile this code via CMake/VS2019. It's sure a whole lot more complicated then it was 25 years ago when I was writing C.
I just did `scoop install cmake`, then built from the command line, was a doddle!
The provided `main.exe` binary worked as-is, but `quantize.exe` did not - I built myself with CMake, and `quantize.exe` started working too.