Google Co-Scientist AI fed previous paper with the answer in it
pivot-to-ai.com
pivot-to-ai.com
> "It's not just that the top hypothesis they provide was the right one," he said.
> "It's that they provide another four, and all of them made sense.
> "And for one of them, we never thought about it, and we're now working on that."
The Google's co-scientist still seems to make a useful assistant.
If an LLM says "my favorite flavor of ice cream is cookie dough", it's just a statistically plausible thing to say.
If I say "my favorite flavor if ice cream is cookie dough", it's because I'm trying to deceive someone, because for some reason I don't want them to know my real favorite flavor is cake batter.
In other words, less thinking and more experiments and observations. Thinking alone leads to ignorance.
As true as "everyone is a calculator". There are superficial similarities, but that we achieve a similar result does not mean that we follow the same reasoning (or lack of).
Finally some logic. Token prediction alone - the behavior - isn't proof of a lack of intelligence.
Better reasoning gives better token predictions. And, sometimes, the ability to invent new tokens entirely.
The trivial interpretation is: every word written can be constructed by optimizing a prediction based on current state, what has been written so far, and a sufficiently complex model. This is true of anything computable: just make the method implicitly contain the program by assigning a high probability to any token that is consistent with running the computation one more step. It's also true of anything expressible: just brute force a solution that can be expressed in n words, then assign a high probability to the first word of these n words.
The profound but wrong interpretation is that intelligence is just statistical prediction according to some general-purpose algorithm, and that this algorithm is tractable. Consider something like solving a SAT problem. You're going to have a hard time using any tractable general-purpose algorithm to predict whether x_2 is true for the satisfying solution based on some long CNF statement plus "x_1 is false".
Now, what you _can_ do is augment your model so that if the previous tokens constitute a CNF-SAT instance plus a partial answer, then you cart these off to a SAT solver and output its next token. But the more you do this, the less force the "mere statistical prediction" part holds. The "next-token predictor" is just an interface to an assembly of different approaches; and often, these approaches (like the SAT solver) will output the whole solution all at once for free.
Animals do have sensors and can reason about real world quite intelligently. This is what LLM- models don't have. Surely, all the text in Internet has a lot of textual information about real world, but something "obvious to animal" might not be there.
I feel us humans are kind of both: We have sensors and experiences about real world, but also oral and written tradition.
Once we develope LLM- with cabability to sense and experience real word (perhaps with some evolutionary algorithm to make it better over time) it should start to be close to human beings.
In other words, I don't think that's a fundamental difference between humans and AI; we just haven't given AI any opportunity to build novel knowledge.
Part of that is because it can't really interact with the world yet, in a significant way. Part of it is because nobody has cracked on-line learning.
Actually when you give AI the ability to interact with a world, as in RL, it does build novel knowledge.
P(tₖ|t₁,..,tₖ₋₁)
For example, the original OpenAI GPT tranformer model was trained to optimise the objective: ∑ log P(tₖ|tₖ₋ₙ,..,tₖ₋₁;Θ)
Where n the length of a sliding "context" window and Θ the parameters of a neural net model (here, a Transformer architecture) used to optimise the objectivbe.See page 3 of the OpenAI paper:
https://cdn.openai.com/research-covers/language-unsupervised...
Also see, for alternative formulations:
Hidden Markov Models:
https://en.wikipedia.org/wiki/Hidden_Markov_model
N-Gram word models:
> P(tₖ|t₁,..,tₖ₋₁)
Not a meaningful definition unless you say what probability distribution it's maximizing it for!
Because for any system that takes any actions (e.g. human beings), there's some probability distribution that that's true for...
Here is the normal distribution probability function.
https://en.wikipedia.org/wiki/Normal_distribution
You can see that there is an actual equation where all the numbers can be filled in except for X, which is the function input.
https://chatgpt.com/share/67bd04cf-1c50-8008-8606-ccac5df989...
And there you have it, novel constructed from available ones, with a genuine pragmatic and LLM specific use case.
https://arxiv.org/abs/2410.05864
Or are you wanting them to use keys that are not on their keyboard?