One thing that's missing from almost all of these energy estimates is the impact of "agentic" queries, which were rare a year ago and increasingly common today.
I've asked questions of GPT-5.6 Pro that take 19 minutes to return an answer, as the model runs hundreds of search and site navigation tool calls (and a bit of executed Python code too) trying to get to a credible answer.
I'm confident that 19 minutes involves a lot more than 75,000 words worth of tokens.
With coding agents we can at least get an accurate token count. I'm already at 204,000 tokens of GPT-5.6 Sol today and I haven't started work yet.