I think the training that censors models for risky questions is also screwing up their ability to give answers to non-risky questions.
I've tried out "Wizard-Vicuna-30B-Uncensored.ggmlv3.q4_K_M.bin" [1] uncensored with just base llama.cpp and it works great. No reluctance to answer any questions. It seems surprisingly good. It seems better than GPT 3.5, but not quite at GPT 4.
Vicuna is way way better than base Llama1 and also Alpaca. I am not completely sure what Wizard adds to it. But it is really good. I've tried a bunch of other models locally, but this one the only one that seemed to truly work.
Given the current performance of Wizard-Vicuna-Uncensored approach with Llama1, I bet it works even better with Llama2.
[1] https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored...