Why are we assuming that shopping agents are going to be using the model that is easiest to fall for prompt injections? Only testing a year old, small model is going to lead to a misleading conclusion.
Why are we assuming that shopping agents are going to be using the model that is easiest to fall for prompt injections? Only testing a year old, small model is going to lead to a misleading conclusion.
Author also needs to make a point, and can’t do so if they are using a powerful model.
next year, astra tier model will be $2/1M or whatever (point is - cheaper) and whatever model is $10/1M will be insane like astra feels today.
also, in any agentic workflow, you use models of various levels together... not just one model of one level...
Model AA Index Cost/task
Claude 4.5 Haiku (thinking) 17 $0.21
GPT-6 Luna (max) 37 $0.07
GPT-6 Sol (low) 34 $0.13
GPT-6 Sol (medium) 40 $0.25
Luna (low) scores higher and costs about 47x less per task and performs better.And Luna (max) scores 37 while still 30% of the cost of Haiku which scored 17.
Why deliberately choose a model that is known to be poor and also costly?
incapable of understanding speculation?
Excessive skepticism does no good.
gpt 6 Luna can be swapped for 5.6 terra or even sol in some cases, and it is much cheaper.
source: i have a lot of evals, and i have ai find the best model for me for those evals over time so i save $ lol