The Opus model that seems to perform better than GPT4 is unfortunately much more expensive than the OpenAI model.
Pricing (input/output per million tokens):
GPT4-turbo: $10/$30
Claude 3 Opus: $15/$75
Pricing (input/output per million tokens):
GPT4-turbo: $10/$30
Claude 3 Opus: $15/$75
My point is that there’s plenty of room for high priced but only slightly better models.
That suggests the inference time is more expensive then the memory needed to load it in the first place I guess?
Probably that and what you mentioned.
Their pricing suggests that either output tokens are more expensive for some technical reason, or they're trying to encourage a specific type of usage pattern, etc.
Nitpick: It's 50% and 150% more respectively.