I kid, I kid.
I kid, I kid.
* Not enough training data - you've used up the Internet (even a percentage of the Internet might be as much as is usable by clever brute force).
* Not enough compute time to fully train (we're not close to that)
* The model covers such a large area that testing is impossible
One thing I'd speculate about is perhaps the more different subjects the program is expected to combine, the more it learns to spout plausible bullshit and clever quips, since for clever humans, that how they relate to stuff they don't know. So "pretentious but uninformed" might be a sign.
I had to rush out the door today after seeing this paper come up so I can’t speak much to its content right now. But if anyone wants to read it and reflect here I’d like to hear it.
I was really impressed by the 002 version. Looking forward to trying out 003 tonight!