It was impressive, but not that much. If you had a basic prompt prefix like:
“The following is a conversation between a human and a helpful AI assistant.
Human: [prompt]\n
AI:”
You would get consistent conversational chat, with the main limit being the model’s limited context window, and obviously frontier intelligence at the time.
I played around quite a bit with davinci-003 and conversational systems before ChatGPT. I kept using this format and API for a while, but at least during the ‘free research preview’ era, found it generally smarter than chatgpt. 3.5-turbo was probably the watershed moment where I moved away from completions on base models; to chat completions.
If you’d like to emulate this experience, spin up a base (not instruction tuned) checkpoint of Llama 1; or a more recent base model. You might underestimate how much ‘intelligence’ you can get :)
Back then the API exposed a lot of controls from sampling to logits, so you could specify a custom stop token (some rarely used Unicode); then switch to near-greedy sampling for more reliable “function calls”, etc; and then switch back to sampling for chat.
In the early days, before releasing something with the API, you had to get your use case approved by OpenAI, like the App Store review process.