Well I was able to run the original code with the 7B model on 16GB vram:
https://news.ycombinator.com/item?id=35013604
The output I got was underwhelming, though I did not attempt any tuning.
The output I got was underwhelming, though I did not attempt any tuning.
https://twitter.com/ggerganov/status/1634310199170179075
I tried it out myself (git pull && make) and the difference in results are day and night! It's amazing to play with, although you should prompt it differently than ChatGPT (more like the GPT-3 API).