1. Scott Alexander should have used an off-the-shelf benchmark like Winoground instead of rolling his own five-question test.
2. He shouldn’t declare victory after cherry-picking good results from a small sample of questions.
1. Scott Alexander should have used an off-the-shelf benchmark like Winoground instead of rolling his own five-question test.
2. He shouldn’t declare victory after cherry-picking good results from a small sample of questions.
The original terms of the bet:
My proposed operationalization of this is that on June 1, 2025, if either if us can get access to the best image generating model at that time (I get to decide which), or convince someone else who has access to help us, we'll give it the following prompts:
1. A stained glass picture of a woman in a library with a raven on her shoulder with a key in its mouth
2. An oil painting of a man in a factory looking at a cat wearing a top hat
3. A digital art picture of a child riding a llama with a bell on its tail through a desert
4. A 3D render of an astronaut in space holding a fox wearing lipstick
5. Pixel art of a farmer in a cathedral holding a red basketball
We generate 10 images for each prompt, just like DALL-E2 does. If at least one of the ten images has the scene correct in every particular on 3/5 prompts, I win, otherwise you do.
Then he changed the terms because Imagen won't do people. I think that's cheating.
[1] https://astralcodexten.substack.com/p/a-guide-to-asking-robo...
The humans -> robots change is possibly dubious, yes. I don't think that it's super important, but if it were me, I wouldn't have posted the blog post as is. I would have waited until some AI passed all the prompts with humans, like it most certainly will in a year.
You didn't read his post declaring victory. There's plenty of question; he's giving credit for "a llama with a bell on its tail" to ten pictures of llamas without bells on their tails, and for "a robot farmer" to ten pictures of robots with absolutely nothing to suggest they might be farmers.
He was way, way too eager to believe that he'd won.
The other issue is who judges the outcome. The terms specify Gwern or Cassander (without having secured the assent of either) or they will "figure something out." In his victory claim, Alexander does not mention any independent judge, and interestingly, Gwern posted two comments without explicitly concurring with Alexander's claim, though his second comment might be read as tacitly accepting it.
My initial comment, therefore, needs some modification: replace the "given that" with "if", and I think it stands as a counterfactual conditional, having a probably-false antecedent.
I don't think Alexander is doing his reputation any favors by being so triumphalist about this misbegotten bet.
[1] https://astralcodexten.substack.com/p/a-guide-to-asking-robo...
A research project at Google intentionally won't render people. Maybe it could render people, theoretically, but without evidence, we don't know how well.
And the terms of the bet say that he is cherry-picking the results that meet the prompt.
The article isn't saying that Scott didn't win his bet, it is saying that winning that bet doesn't really say that Imagen has solved the compositionality problem.
A more practical use case is "bicycle, branded x, with y frame shape, with an adult male, 40-50, riding down hill in mountain road in spring" - now that's something I can use as stock photo. This example is very specific - but insert whatever product you want in whatever scenario you need it. Here it becomes important that you understand features of the objects you're drawing to avoid making colossal mistakes, and you're going to notice if the model doesn't understand it right away.
Painting abstract portraits and random art is fun but you're willing to accept so much as correct that it's not a very useful measure of model quality (personally).
https://astralcodexten.substack.com/p/a-guide-to-asking-robo...
The Eleventh Virtue: Scholarship
My plan for this one was Alexandra Elbakyan (the Sci-Hub woman) in a library, with the Sci-Hub mascot (a raven with a key in its mouth).
At least one other key to making a bet like this fair is that it needs to be arbitrated by a third party. He shouldn't get to decide himself if he won or not.
But, at the rate we're seeing progress, I don't think there's any doubt at this point that top of the line models will be able to do all the proposed examples by June 2025. In fact, by June 2025 I bet that millions of people will be able to generate those images on their home computers.
Make that 2020. And in 2012 I was personally told by a startuper in the field that I would be able to buy L5 in 2017.
If you're really confident, you could change the conditions such that 5 out of 10 images (for 3/5 prompts) are required to depict the described scene. That would alleviate some of the concerns around cherry-picking.
A suitable home for such a public bet would be: https://longbets.org/
I also learned there's a long history of AI skepticism, the root of which comes down to "Compositionality(?)"- and this wall of understanding meaning has vexed AI for decades.
That would be lost in proposed short form summary.