Can somebody please clarify. Is the cost $0.002 per 1k tokens generated, read, or both?
This pricing model seems fair since you can pass in huge prompts and request a single word reply, or a few words that expect a large reply
Such an approach could in theory make it so you spend a little upfront to train more complex (read: concepts costing many tokens) and can subsequently reuse it cheaply because you're using an embedding of the vectors for that complex concept instead which may only take a single token.
gpt-3.5-turbo-0301, 2 requests 28 prompt + 64 completion = 92 tokens