> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful?
Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that:
A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU. There, inference costs declined from $15 per million tokens in May 2024 to $0.12 per million tokens by December 2024 (Phi 4).
- https://hai.stanford.edu/assets/files/hai_ai-index-report-20...15/0.12 -> factor of 125 cost reduction in 7 months.
But that may well be an extreme case. To show how broad the range is, another quote from the same publication:
Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.