Also the author is deliberately chooses a model that is 1 year old and was supposed to be underpowered even then. If cost was deciding factor, they could’ve chosen Luna.
Model AA Index Cost/task
Claude 4.5 Haiku (thinking) 17 $0.21
GPT-6 Luna (max) 37 $0.07
GPT-6 Sol (low) 34 $0.13
GPT-6 Sol (medium) 40 $0.25
Luna (low) scores higher and costs about 47x less per task and performs better.And Luna (max) scores 37 while still 30% of the cost of Haiku which scored 17.
Why deliberately choose a model that is known to be poor and also costly?