It seems sensible to extract features and reason about things the way a human would, but it turns out its easier to scale pattern matching purely done by computer.
It seems sensible to extract features and reason about things the way a human would, but it turns out its easier to scale pattern matching purely done by computer.
https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done."
My point being, optimization or splitting up int subs, before handing over the problem to the machine, makes sense.
- We don't have limitless CPU cycles
- Thus we need to split things into sub-problems
If so that might still be amenable to the bitter lesson, where Sutton is saying human heuristics will always lose out to computational methods at scale.
Meaning something like:
- We split up the thought to vision problem into N sub-problems based on some heuristic.
- We develop a method which works with our CPU cycle constraint (it isn't some probe -> CPU interface). Perhaps it uses our voice or something as a proxy for our thoughts, and some composition of models.
Sutton would say:
Yeah that's fine, but if we had the limitless CPU cycles/adequate technology, the solution of probe -> CPU would be better than what we develop.
But i think we're onto something!
Voice to image indeed might give better results than text to image, since voice has some vibe to it (intonation, tone, color, stress on certain words, speed and probably even traits we don't know yet) that will color or even drastically influence the image output.
With image generation on the other hand, which starts from a handful of words, we can first do some text processing into categories, such as objects vs people, color vs brightness, environment vs main object, etc.