> the harder it is for me to use these tools in a way that doesn’t feel like too much blind faith (even if it works!)
I tend to ask multiple models and if they all give me roughly the same answer, then it's probably right.
I tend to ask multiple models and if they all give me roughly the same answer, then it's probably right.
You can see this effect in the ARC-AGI evals, too much context impacts even o3(high).
... or they had a lot of overlapping training data in that area.