ie. you would fail too.
ie. you would fail too.
Anyways, that's besides my point. The point of mine is that, it always turn into all-caps flamewars like this, with no middle ground or third camps, and that this has to be more of a phenomenon than regular disagreements. This isn't bikeshedding. This is Spanish bullfighting centered around a piece of red cloth.
1: https://news.ycombinator.com/item?id=42216694
2: I just asked Gemini "is 60% accuracy over 11k participant for a test statistically significant and why", it said "yes, it is overwhelmingly statistically significant" and "completely off the charts". They said p<0.05 figure would be 50.94%.
Also I'm not that worried about the adversarial conditions, any real life conditions are likely adversarial in the relevant sense. No one is one-shotting generation and serving you that, obviously output has been selected for quality. I would call Scott's test not adversarial, but fair.