OpenHermes-2.5-Mistral-7B is better than GPT-3 (and scores even better than GPT-3.5-Turbo in human evaluations) and can even run on a raspberry pi or in the browser. On a laptop CPU it uses about 5GB of RAM (in 5bit) and runs around 20-30 tokens per second, which is very fast.
I recommend downloading and running OpenHermes inside LM Studio. https://lmstudio.ai/
In LM Studio, search for OpenHermes. Pick the Q5_K_M version (this is the best quality/speed trade off). Then go to the chat tab.
On the chat tab, set the context length to 4096 (or up to 16k if you want longer context) and set the number of CPU cores you have under "Hardware Settings."
Select the model from the drop down and start chatting!