From my experimentation I suspect there's some subtle bug in llama.cpp that especially degrades code related prompts- even without quantizing
I've mostly heard that, at least for the larger models, quantization has barely any noticeable effect. Would be nice to witness it myself.
If you're not trying to get it to be a chatbot it's much easier, here's a prompt that worked for me on the first try in the default mode with 13Bq4 on a 1080Ti:
Here are is a short, clear, well written example of a program that lists the first 10 numbers of the fibonacci sequence, written in javascript:
```js
and when given that it finished it with: function Fib(n) {
if (n == 0 || n == 1) return 1;
else return Fib(n-1)+Fib(n-2);
}
var i = 0;
while (i < 10) {
console.log("The number " + i + " is: " + Fib(i));
i++;
}
```
\end{code}So under the hoods, ChatGPT is just a model like Llama where they prepend every user input with a context that makes it behave like a chatbot?
Thanks for the info.
65B fp16 is ungodly slow, ~300,000 ms/token on the same machine.