Brilliant, I have always felt that one of the major problems with machine learning, consequently LLMs, is the boring average based loss functions that under-represent the unique and the rare. It seems our collective civilization is using a similar function and heading in the same direction of optimizing for the average.