how many times do you have to be metaphorically hit in the head with a brick before you realize inference margins at api pricing were 80%+
1. The basic mechanism is literally described in the post: they found efficiencies and passed it down
2. This has been the trend for all the time LLMs have existed
Why confusion then? What’s surprising you?