Let's not do this again, please
bloodinthemachine.com
bloodinthemachine.com
I think this is untrue. Sora is doing way more than stitching existing videos and images together.
The relentless underhyping of AI is almost as annoying as the overhyping. ChatGPT, for example, is a non-human entity that speaks fluent English. This is an amazing achievement, no matter how many people try to downplay it.
[edit]
I'm not sure what i would call its actions, if not speech though. It does not think independently, but it does pass information. I guess it computes an answer.
I think several LLMs have sufficiently demonstrated the ability to communicate concepts using natural language to be able to say that the statement is true.
If you tie an LLM to underlying software (and thus also ultimately hardware, if desired), it can take your instructions in natural language, translate them into a form suitable for the software to process, and then take the output and render that back into intelligible natural language.
On that basis, I would argue that the statement "it speaks fluent English" is essentially correct.
If you assert "but that has nothing to do with intelligence", you may or may not be correct, but -combining personal empirical observation with your world model- , intelligence would then appear to be orthogonal to the ability to speak English.
Either that, or there is a flaw in your world model.
Whatever the case may be, LLMs are quite evidently capable of communicating in English.
[Edit]
And the dream isn't one of mine, the parent references "sci fi dream."
Are we still arguing whether "able to speak english" is or is not a subset of "able to communicate in natural language", or are we in agreement in our entirety now?
And there are many "sci fi dream"s of course, one of which is/was for computers to be able to communicate in natural language. (See eg: star trek, iron man)
You clearly have a point you'd like to make, but I don't think claiming that ChatGPT can't "speak" English is a effective way to make it. It can.
Someone mentioned in fine arts they make a distinction between craft and art, which is an excellent point.
[Edit]
https://news.ycombinator.com/item?id=39423510
[Edit]
If all it takes to "speak" is vocalizing sounds, does text to speech count?
[Edit]
I guess it's in the name, text to "speech." Again, gpts do not have independent agency.
GPT responds in a "series of plausible tokens" that would pass as fluent English if generated by a human.
Let's put aside the vocalization point, that's a distraction.
while True:
print('a')
GPT-n (not chat) generates continuations to text in which the probability distribution for the next output it produces matches the distribution of outputs after similar text in the training corpus, by most metrics you might examine. Trivial examples of those metrics include word and n-gram frequencies, but also it turns out that at sufficiently low loss that looks like "given a couple of input texts and a context which affords insightful commentary, produce insightful commentary at the base rate".There are, of course, caveats. Notably:
- ChatGPT has its own stuff going on around chat tuning, tool use, tuning to be less likely to produce outputs OpenAI doesn't want, etc
- The statistical regularity thing is not magic. If it requires more than n_layers steps of computation to determine the sensible next output, the model will not be able to do that. I think the canonical example here is usually having the model complete something like "'d072c916029965a7676da4244160c413e31bc8a0' == sha1('I saw it on hn'); '265149165fcb742a900a44b8f123885dc6ac5d12' == sha1('" and then having the model brute-force sha1 -- obviously it's not going to be able to do that.
- The model generates text sequentially, one token at a time. This is importantly not the process by which most text in the training corpus was written. So in the cases where earlier text importantly depends on text which was written earlier temporally, but which occurs later in the string, the model will be likely to make mistakes (where a "mistake" is "writing text which is statistically surprising, relative to the training corpus").
I just feel like we're not in a position to even begin understanding our disagreement until you at least recognize the question or point I'm trying to get here. If you disagree with the question, or don't understand it: why?
But I don't think there's anything going on in language generation beyond "based on the current context, produce an appropriate output token for that context based on the observed and inferred distribution of training inputs". I think that's also how human language generation works, though "the context" for humans includes a lot more than just a few thousand words of text. But I think that the surprising thing about e.g. GPT-4 is how well it does the thing, rather than the fact that it does the thing at all.
"Computers can't actually do calculations, they're just a bunch of circuits"
"Cars don't actually move, they're just a heat engine"
"Chat GPT doesn't speak and understand natural language, it just performs statistical analysis"
-
Eppur si muove.
-
We define cells as being alive.
We use computers to do maths.
Cars are used to travel from A to B.
And an LLM can accept instructions in a natural language, carry them out, and report back in natural language.
[Edit]
Alive is also a word with a lot of ambiguity associated. I'm not sure we have a clear definition. If "I think therefore I am" is the requirement, gpts do not qualify, as they do not act independently.
[Edit]
Definition 1) To produce words by means of sounds; talk. - gpt passes
Definition 2) To express thoughts or feelings to convey information in speech or writing. - gpt fails
[Edit]
You used quotes, but I never said most of those things.
Definition 3) To convey information or ideas in text. (gpt passes this one)
By the way, dictionary definitions are often very fuzzy, so this is not really any sort of great evidence for anything. More me returning your lackadaisical volley with the same vim and vigor as I received it. ;-)
I bet you can do better!
I didn't bother to include the third definition as I'm not sure how it is different than either of the first two.
So what if it passes the third?
And: selectively leaving out part of a quote that appears to (partially) contradict your thesis might be seen by some as an attempt to mislead. Best to err on the side of caution!
Though -like I said- dictionaries are possibly the weakest source of definitions you can find. Dictionaries need to keep things extremely short, so they're unlikely to have much nuance.
Even Encyclopedias (including Wikipedia) are more accurate and reliable, simply because they have more space to expand on a concept in detail.
Books or papers are best.
[1] https://news.ycombinator.com/newsguidelines.html "Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith. "
[Edit]
i.e. "To produce words by means of sounds or text." Whoops.
Anything less than Midjourney for video will be hugely disappointing.
I still love chatGPT but my use is so limited compared to what it was 9 months ago. Then I think about how a language model could help with basically anything I have ever been paid to do and the reality is I don't think it would be much help. Then if we go back to college/school I wouldn't have used it to learn more. I would have used it to do less work and probably learned less.
If someone spent countless hours crafting prompts to generate a really cool image or something, then that's much less impressive. It's kind of like the sales demo vs. the actual product you receive.
Transcripting and summarizing meetings was never that easy.
It is already a game changing technology..
I'm lost on how you dismiss this.
Crypto on the other hand has not solved any problem and most peo still don't get it or even used it.
And very often we get new great ai things.
Nearly my whole family tried chatgpt.
Replacing human labor at scale will take a generation to work out and this is day one. They'll be happy to go through pain points like this for the next decade in order to learn how to make it work.
It's an impressive demo which got viral for much better reasons than a lot of other things before.
And I really don't mind at all if super rich people invest with even more money into it.
Surely this is a typo - it should be “rarely” had a chance to ask.