But to be honest, what surprised me most about that was that Jane Street cares about counting the token usage in its API calls. I can see why they would care if they were interacting with AI models at very large scales, but are they? They have billions of dollars, so it struck me as odd to see such concern for efficiency in what must be a very small cost center.
I wonder if the motivation is more about providing feedback to the users of the frontend tool, as they write their prompts, so they can gain an intuition for what makes an efficiently tokenized prompt? (So that when the day comes that they are running these prompts at scale, the token sizes really matter). Or maybe I'm imagining "prompts" wrong, and these are not just a few paragraphs of text, but actually 20,000+ tokens at a time, including data like e.g. pasted CSVs of trading history for the LLM to analyze?