It's the same as asking a person to double check; it works because the models know different things. The next step would be to use a lightweight model to automate the ensembling...
It's the same as asking a person to double check; it works because the models know different things. The next step would be to use a lightweight model to automate the ensembling...
I think double checking is better than not, but without the ability to really “know” reason it feels a bit like adding one more hull layer to the titanic in an effort to make it unsinkable.
How do you tell if it's actually stating the reasoning that got it to its answer originally, as opposed to constructing a plausible-sounding explanation after the fact? Or is the goal just to see if it detects mistakes, rather than to actually get it to explain how it arrived at the answer?
But yes, it'll be more effective on a different model.
Do you find you generally get a well reasoned outcome or do you also find the model stretching to come up with a take that aligns with your skepticism?