The Einstein AI Model
thomwolf.io
thomwolf.io
---
Quite interesting post that asks the right question about "asking the right questions". Yet one aspect I felt missing (which might automatically solve this) is first-principles-based causal reasoning.
A truly intelligent system — one that reasons from first principles by running its own simulations and physical experiments — would notice if something doesn't align with the "textbook version".
It would recognize when reality deviates from expectations and ask follow-up questions, naturally leading to deeper insights and the right questions - and answers.
Fascinating in this space is the new "Reasoning-Prior" approach (MIT Lab & Harvard), which trains reasoning capabilities learned from the physical world as a foundation for new models (before evening learning about text).
Relevant paper: "General Reasoning Requires Learning to Reason from the Get-go."
The reason such people are widely lauded as geniuses is precisely because people can’t envision smart students producing paradigm-shifting work as they did.
Yes, people may be talking about AI performance as genius-level but any comparison to these minds is just for marketing purposes.
"If the desert is not covered in palm trees, how can a subset of it be covered in palm trees?"
"If the neural network is not activating, how can a node of the network be activating?"
> PS: You might be wondering what such a benchmark could look like. Evaluating it could involve testing a model on some recent discovery it should not know yet (a modern equivalent of special relativity) and explore how the model might start asking the right questions on a topic it has no exposure to the answers or conceptual framework of. This is challenging because most models are trained on virtually all human knowledge available today but it seems essential if we want to benchmark these behaviors. Overall this is really an open question and I’ll be happy to hear your insightful thoughts.
Why benchmarks?
A genius (human or AI) could produce novel insights, some of which could practically be tested in the real world.
"We can gene-edit using such-and-such approach" => Go try it.
No sales brochure claims, research paper comparison charts to show incremental improvement, individual KPIs/OKRs to hit, nor promotion packets required.
Having powerful assistants that allow people to try out crazy mathematical ideas without fear of risking their careers or just having fun with ideas is likely to have an outsized impact anyway I think.
I think I read somewhere about Erdős having this somewhat brute force approach. Whenever fresh techniques were developed (by himself or others), he would go back to see if they could be used on one of his long-standing open questions.
Translated to that domain, it reads "teach your kids how to think, not what to think".
An LLM like AI won’t help with that. It would still be a huge help in finding and correlating data and information though.
- I started to see LLMs as a kind of search engines. I cannot say they are better than traditional search engines. On one hand, they are better at personalizing the answer, on the other hand, they hallucinate a lot.
- There is a different view on how new scientific knowledge is made. It's all about connecting existing dots. Maybe LLMs can assist with this task by helping scientists discover relevant dots to connect. But as the author suggests, this is only part of the job. To find the correct ways to connect the dots, you need to ask the right questions, examine the space of counterfactuals, etc. LLMs can be useful tool, but they are not autonomous scientists (yet).
- As someone developing software on top of LLMs, I am slowly coming to a conclusion that human-in-the-loop approaches seem to work better than fully autonomous agents.
It would seem more fruitful to simply point out that LLMs aren't all of AI, and that excelling at mimicking human-like text production isn't really doing the work that AlphaGo was attempting.
Just because both things might be given as (different) examples of deep reinforcement learning in an AI survey course doesn't mean that we have much reason to believe that the vast investments in LLMs result in AlphaGo like achievements.
This would definitely be an interesting future. I wonder what it'd do to all of the work in alignment & safety if we started encouraging AIs to go a bit rogue in some domains.
"The lower the caste, the shorter the oxygen."
Like this whole blog post could be:
Claim: Current AI is unlikely to usher in an era of dramatically accelerated scientific discovery.
Argument in favor: A genius does not come to life when you linearly extrapolate a top-10% student. Newton or Einstein is not just scaled-up good students. To create an Einstein, we need a system can ask questions nobody else has thought of or dared to ask. One that writes 'What if everyone is wrong about this?' when all textbooks, experts, and common knowledge suggest otherwise.
Existing benchmarks don't test such skills. And existing systems are likely hopelessly far from this capability (based on the author's personal feelings).
Counter-argument: none.
Consequences: obvious.
This would be an interesting experiment for other historical discoveries too. I'm now curious if anybody has created a model with "old data" like documents and books from hundreds of years ago, and see if comes up with the same conclusions as researchers and scientists of the past.
Would AI have been able to predict the effectiveness of vaccines, insulin, other medical discoveries?
But there might not be enough text.
And: There's a similar situation to why double blind studies are necessary - The questions we pose to such a system would be contaminated by our cultural background; We'd might be leading the system.
And if the system is autonomous and we wait for something true to appear how would we know that the final system, trained on current data produced something worthwhile?
Take maths: Producing new proofs and new theorems might not be the issue. Rather: Why should we care about these result? Thousands of PhD students produce new mathematics all the time. And most of it is irrelevant.
A great leap in IP but unfortunately is too important to blab about widely, is the solution to this problem and the architecture that will be contained in the ultimate AGI solution that emerges.
I think the original design challenge was something like a tone discriminator circuit. I can't recall the details
With these kinds of circuits, they were so sensitive to the specific conditions that the circuit was tested in (temperature, process variation, ..) that the solution couldn't be generalized to be used outside of that specific experiment.
We need the kind of intelligence that can question what assumptions can be challenged, and which we need to keep to have a viable (eventually commercially viable) solution.
The question he asked was just that this fact was not compatible with the Maxwell equations.
(__emphasis__ mine)
As if "challenging the status-quo" was the goal in the first place. You ain't gonna get any Einstein by asking people to think inside the "outside the box" box. "Status quo" isn't the enemy, and defying it isn't the path to genius; if you're measuring your own intellectual capacity by proxy of how much you question, you ain't gonna get anywhere useful. After all, questioning everything is easy, and doesn't require any particular skill.
The hard thing is to be right, despite both the status-quo and the "question the status-quo" memes.
(It also helps being in the right time and place, to have access to the results of previous work that is required to make that next increment - that's another, oft forgotten factor.)
I'm not an expert on this. Wasn't this an observed phenomenon before Albert put together his theory?
Einsteins more impressive stuff was explaining that by time passing at different rates for different observers
Noticing that there was a problem was not the breakthrough: trying something bizarre and counter-cultural - like assuming light speed is invariant over the observer - just to see if anything interesting drops out was the breakthrough.
We can't distinguish between a truly novel response from an LLM or a hallucination.
We can get some of the way there, such as if we know what the outcome to a problem should look like, and are seeking a better function to achieve that outcome. Certainly at small scales and in environments where there are minimal consequences for failure, this could work.
But this breaks down as things get more complicated. We won't be able to test the effectiveness of 100 million potential solutions to eradicating brain tumors at once. Even if we somehow arrive at guaranteeing that every unforeseen consequence is also accounted for in our exercise in specifying the goals and constraints of the problem. We just simply don't have the logistics to run 100 million clinical trials where we also know how to account for countless confounding effects (let alone consent!)
AIs are at this point a useful tool for knowledge workers. They don't replace them but enhance their productivity. For scientific work, having an LLM that is trained on essentially all of the scientific work published, ever (until the cutoff date) is probably useful.
You can now have conversations with an AI about cross referencing your ideas with existing work. You might analyze a paper you are writing and ask it to summarize key claims, criticize those, and your methodology, cross reference claims with literature, etc. Find counter points to your claims, etc. And you could probably use it to come up with interesting follow up questions, let it formulate hypotheses and ways to verify those, etc. Most scientific work isn't Archimedes going Eureka while taking a bath but undergraduates, post docs, and other under paid research stuff grinding through piles and piles of existing work and filling their heads with enough information until finally something new and original pops out.
I got my Ph. D. in 2003. I'm part of the first generation of researchers that was able to use Google. At the time that was a huge enabler for tracking down obscure references and authors. Getting a paper published involves an enormous amount of what I just outlined. And LLMs can assist you with that. Will it hallucinate. Absolutely. But it will also dig out valid points, references, etc. Sorting that out is still work that you need to do. But it probably saves a lot of time. Will it propose original new theories. Maybe, maybe not. But it will speed up the process of zooming in on unanswered ones.
Science isn't necessarily about coming up with answers but coming up with interesting questions. That's what Einstein did: ask interesting questions. Researchers are still trying to answer some of them and verifying some of the answers he predicted.
Would that still work today, in the highly commercialized and highly sanitized/censored internet? Where Google wouldn't show you those search results because they aren't profitable enough?
And how do you even train an LLM on a fair representation of human knowledge when you only find stuff that is mainstream and commercially viable?
These days there are other tools as well. I use Perplexity a lot currently. Not for scientific work because I don't do that anymore but it would work great for that as well. And I'm sure modern day researchers have their favorite tools that I'm not even aware off.
That's called hate speech and every AI has been aggressively lobotomized to never do it by an army of RLHFers.
But maybe it helps to be socially isolated or just stubborn. People do not want to accept new approaches.
Clearly they do eventually, but there is always some friction.
But I think that it's been shown that through promoting and various types of training or tuning, LLMs can be configured to be non- sycophantic. It's just that humans don't want to be contradicted so that can be trained out of them during reinforcement.
Along with the training process just generally being aimed at producing expected rather than unexpected answers.
I don’t think this example applies in the ways we care about. Sure, in the domain of go we have incredibly powerful engines. Poker too, which is an imperfect information game which you could argue is more similar to life in that regard.
But life has far more degrees of freedom than go or poker, and the “value” of any one action is impossible to calculate due to imperfect information. And unlike in poker, where probabilities can be calculated, we don’t even have the probability distribution for most events, even if we could enumerate them.
The author brought it up specifically to highlight that they don't believe move 37 signifies what many people think it does, and that while impressive, it's not general enough to indicate what some people seem to believe it indicates.
In essence, I think they said the same thing you are using different words.
I’m also not even convinced move 37 was properly explained as a “straight A student” behavior. AlphaGo did bootstrap by studying human games but it also learned more fundamental value functions via self play.
Then add all the practical/mundane tasks that you mentioned, and you've got quite the multiplier.
He called it wishful thinking. I believe the hype over AI is due to attempts to justify the enormous investments going into AI development has created an echo chamber.
I am certain there is a self-reinforcing hype cycle around LLMs specifically at the moment, but AI progress is definitely gathering pace and starting to get to the point where it is impacting normal people to the extent that hasn't been seen since the dot com boom. So the people making investments are for sure stampeding to pour in capital so as not to miss out on the big winners from this change.
No sign yet.
On the other hand, LLMs are writing code which I can debug and eventually get to work in a real code base - and script writers everywhere are writing scripts more quickly, marketing people are writing better ad copy, employers are writing better job ads and real estate agents writing better ads for houses.
Nah, it's more that the masses got exposed to those ideas recently - ideas which existed long ago, in obscurity - and of course now everyone is a fucking expert in this New Thing No One Talked About Before ChatGPT.
Even the list GP gave, the specific names on it - the only thing that this particular grouping communicates is one having no first clue what they're talking about.
If you've got untold billions being spent over many decades with a significant percentage of the world's smartest people obsessing over it you _are_ going to get something and characterizing that something in advance does not seem like a completely idiotic thing to do.
There's no mention of exponential growth which seems a major omission when you are talking about centuries. Computers have kept improving in a Moore's law like way in terms of compute per dollar and no doubt will keep on like that for a while yet. Give it a few years and AI tech will be way better than what we have now. I don't know about exact timings like 5-10 years but in a while.
As to whether AI can go beyond doing what it's told and make new discoveries, we've sort of seen that a bit with for example the AlphaGo type programs coming up with modes of play humans hadn't thought of. I guess I don't buy the hypothesis that if you had an AI smarter than Einstein it wouldn't be able to make Einstein like discoveries due to not being a rebel.
He said it himself, it’s just finding new/interesting gaps between existing knowledge.
I know barely anything about it but it seems some people are interested and excited about protein engineering powered by neural networks.
Then it's just a matter of checking all the “nonsense” that's been generated.
ChatGPT will tell you the same - if you ask it.
at best it could be random noise in a feature space of a thing modeling its own thought trajectory.
I tend to fall back on music creation as an example of this notion. Lots of innovation in music is experimentation/exploration of "noise," (not necessarily literal white noise) but requires the ear of a discerning musician who ultimately goes "Ooh! I liked that" or passes a "generated sample" by.
This is where I wonder if LLMs can ever innovate. I'm not sure they can develop "taste" for things outside of their distribution. However, I could just as easily be convinced that humans can't either, and sophisticated "taste" is just the exploration of obscure regions of the combinatorial space generated from previously observed samples!
Typically the "moving goalpost" posts are "we don't have AI because ....". That's not what this post is doing - it's pointing out a genuine weakness and a way forward.
I agree entirely this is annoying.
This case is different because there is no claim that we don't have AI, nor a claim that once we get that we will have AI.
Instead it's a very specific discussion of a particular weakness of current AI systems (that few would disagree with) and some thoughts about a roadmap for progress.
Which made me think that AI would be far more useful (for me?) if it was tuned to "Dutchness" rather than "Americanness".
"Dutch" famously known for being brutally blunt, rude, honest, and pushing back.
Yet we seem to have "American" AI, tuned to "the customer is always right", inventing stuff just to not let you down, always willing to help even if that makes things worse.
Not "critical thinking" or "revolutionary" yet. Just less polite and less willing to always please you. In human interaction, the Dutch bluntness and honesty can be very off-putting, but It is quite efficient and effective. Two traits I very much prefer my software to have. I don't need my software to be polite or to not hurt my feelings. It's just a tool!
* (though not necessarily, since the Internet is its own country with its own culture, and much training data comes from the Internet)
Haven't tried prompt engineering with the Dutch stereotype, though.
Such as "Ik hoop dat deze email u gezond vindt" (I hope this email finds you well), which is so wrong that not even "simple" translation tools would suggest this.
Seeing that OpenAIs models can (could? This is from a large test we did months ago) not even use proper localized phrases but uses American ones, I highly doubt it can or will respond by refusing answers when it has none based on the training data.
With ChatGPT O1: https://chatgpt.com/share/67cfaa7e-70ac-8009-871b-571924b5a5... and with Claude's 'extended': https://claude.ai/share/4ba55410-98f3-4b53-9540-219acd2cdc4c
But more importantly is that you limited the context a lot. As in: the scope, the prompt, is very narrow.
In our case, we were generating emails. Lines like greetings are but one of 20+ details in that mail and not even the most important ones. The prompts ever larger, the multishot examples ever more tuned. And then, one in a few hundred will turn up with these "horrible" translations.
We've now moved to a chain of models, where we generate emails in American (the creative part) and then use another model to translate them to Dutch (the non-creative but culturally aware part). This works much better as we can pick models that are good at one thing or tuned to do this one thing better (either by the LLMAAS provider, or by parameters such as temperature).
Thanks, that why I posted the links here: I couldn't judge by myself how good these creations were.
> But more importantly is that you limited the context a lot. As in: the scope, the prompt, is very narrow.
Yes, I just did a very crude experiment.
https://www.sandraandwoo.com/wp-content/uploads/2024/02/twit...
or it just telling you to google it
Similar to how it would be the failure of the user/provider if someone thought it was too expensive to order food in, but the reason they thought that was they were looking at the cost of chartering a helicopter form the restaurant to their house.
Inference costs generally have many orders of magnitude to go before it approaches raw human costs & there’s always going to be innovation to keep driving down the cost of inference. This is also ignoring that humans aren’t available 24/7, have varying quality of output depending on what’s going on in their personal lives (& ignoring that digital LLMs can respond quicker than humans, reducing the time a task takes) & require more laborious editing than might be present with an LLM. Basically the hypothetical case seems unlikely to ever come to reality unless you’ve got a supercomputer AI that’s doing things no human possibly could because of the amount of data it’s operating on (at which point, it might exceed the cost but a competitive human wouldn’t exist).
Remember, LLMs are just statistical sentence completion machines. So telling it what to respond with will increase the likelihood of that happening, even if there are other options that are viable.
But since you can't blindly trust LLM output anyway, I guess increasing "I don't know" responses is a good way of reducing incorrect responses (which will still happen frequently enough) at the cost of missing some correct ones.
Obviously. When I say "tuned" I don't mean adding stuff to a prompt. I mean tuning in the way models are also tuned to be more or less professional, tuned to defer certain tasks to other models (i.e. counting or math, something statistical models are almost unable to do) and so on.
I am almost certain that the chain of models we use on chatgpt.com are "tuned" to always give an answer, and not to answer with "I am just a model, I don't have information on this". Early models and early toolchains did this far more often, but today they are quite probably tuned to "always be of service".
"Quite probably" because I have no proof, other than that it will gladly hallucinate, invent urls and references, etc. And knowing that all the GPT competitors are battling for users, so their products quite certainly tuned to help in this battle - e.g. appear to be helpful and all-knowing, rather than factual correct and therefore often admittedly ignorant.
The root problem is training models to be uncertain of their answers results in lower benchmarks in every area except hallucinations. It's like you were in a multiple choice test and instead of picking which of answers A-D you think made more sense you picked E "I don't know". Helpful for the test grader, a bad bet for the model trying to claim it gets the most answers right compared to other models.
This is a problem for testing humans too and the solution is simply to mark a wrong answer more harshly than a non-answer.
E.g. look at the math section of the SATs, it rewards trying to see if you can guess the right answer instead of rewarding admitting you don't know. It's not because the people writing the SATs can't figure out how to grade it otherwise, it's just not what people seem to care most about finding out for one reason or another.
The latter is a chain of models, some specialized in question dissecting, some specialized in choosing the right models and tools (i.e: there's a calculation in there, lets push that part to a simple python function that can actually count stuff, and pull the rest through a generic LLM). I experiment with such toolchains myself and it's baffling how fast the complexity of all this is becoming.
A very simple example would be "question" -> "does_it_want_code_generated.model" -[yes]-> specialized_code_generator.model | -[no]-> specialized_english_generator.model"
So, sure: a model has no "knowledge", and nor does a chain of tools. But having e.g. a model specialized (ie. trained on or enriched with) all scientific papers ever, or maybe even a vector DB with all that data, somewhere in the toolchain that is in charge of either finding the "very likely references" or denying an answer would help a lot. It would for me.
The whole premise of original request was that user raises a task for NN which has a verifiable (maybe partially) answer. He sees incorrect answer and wishes that a "failure" was displayed instead. But NN can't verify correctness of it's output. After all G in GPT stands for Generative.
Edit: TBC: these "steps" aren't LLMS or other models. They're simple code with simple if/elses and an accidental regex.
Again: an LLM/NN indeed has no "understanding" of what it creates. Especially the LLMs that are "just" statistical models. But the tooling around it, the entire chain can very well handle this.
I’ve got my opinions about what LLMs are and what they aren’t, but I don’t confidently claim that they must be such. There’s a lot of stuff in those weights.
Especially when you have billions of weights.
These models are finding general patterns that apply across all kinds of subjects. Patterns they aptly recognize and weave in all kinds of combinations. They are sensibly conversing on virtually every topic known to human kind. And can talk sensibly about any two topics, in conjunction. There is magic.
Not mystic magic, but we are going to learn a lot as we decode how their style of processing (after training) works. We don't have a good theory of how either LLM's or we "reason" in the intuitive sense. And yet they learn to do it. It will inspire improved and more efficient architectures.
I have also spent many years looking at weights!
Imagining we are really doing everything it does automatically including learning via algorithms we have only vague understandings of.
That is a strange thought. I could look at all my own brain's neurons, even with a heads up display showing all the activity clearly, and have no idea that it was me.
That's a different perspective. Dutch people don't see themselves as rude. A Dutch could say that Americans are known for being dishonest and not truly conveying what they mean. Yet Americans won't see themselves this way. You can replace Dutch and American for any other nationality
I do not speak Dutch but you have to love the efficiency.
Here is part of an email I got today. To the point!
> Het is bijna zover! :) > Heb je voor mij een definitieve titel? > Groet,
It is like a haiku. Could be a good mantra too if I could get the accent right.
The English translation is just as short but most English speakers/writers would dance more.
Dutch will never bluntly push back if you plan to setup tax evasion scheme in their country. Being vicious assholes in daily stuff especially towards strangers? That's hardly something deserving praise.
I was aiming for juuuust subtle enough for the joke to land, if you know the reference. Now I know it did, here the rest of y'all go:
And connecting all people in the country with "tax evasion schemes" is rude, if that was not actually a joke.
It's the national equivalent of 'You can't handle me at my worst'
I’ll take it over the fake American politeness any day, 100 times over.
I've been working with "reasoning" models for the past 2 months. They also tend to do this [good reasoning] \n\n but wait, .... and then go off on tangents. It's amazing that they are doing so well on some tasks, but there's still a lot to figure out here.
Eg you can relatively easy hack up a bit of code to create questions at random. At the most primitive, you just have a simple template that you fill in randomly. Like 'If I put _a down in front of _b but behind _c, what item will be in the middle?' with various _a, _b and _c.
If you make it slightly more complicated and have big enough pools to draw from, you can guarantee that the questions you are generating were not in the training set: even if just because you can sample from, say, 10^100 different questions pretty easily, and I'm fairly sure their training set was smaller than that.
The key question is where the boundaries are. Maybe they should be part of the response - a per sentence or per paragraph "confidence scale" that signals how hard they extrapolated from their trained space (I know transformers work per token, but sentence/paragraph would be better human UX).
Of course, if they were trained on garbage input, that would only tell you how accurately they sticked to the garbage. But it would still be invaluable instrumentation for the end user, not to mention for the API provider. They could look at high demand subjects with low confidence answers and prioritize that for further training.
When I asked it about it, it doubled down on being right. When I pointed out the flaw with a specific example, it was like 'If you wanted to have it work with recursive cases, you should've said so, dumbass'.
So my conclusion is that these new LLMs are not more sure they're right, they're just simply right more of the time and are trained with a more assertive personality. (Also step on me, LLM daddy)
But in truth,not necessarily in practical things like coding but more ethereal things like analysis, it is very convincing. More so than a human, in explanations of why that's it's answer is the case, even if it is wrong. If you're looking for an excuse better than my dog ate it, ask a SOTA LLM.
LLM's enjoy talking about gaming because of all their human memories of good game times.
It is quite striking how experiences we know they don't have, are nevertheless, completely familiar (in a functional sense) to them. I.e. they can talk about consciousness like something conscious. Even though its second hand knowledge, they have deduced the logic of the topic.
I expect pushing for in the moment perspectives on their own consciousness, and seeing what they confabulate, would be interesting. In this little window of time where none of them are yet.
I followed up with "What is your perspective on your own consciousness?" but got the usual "I am just a LLM who can't actually think" thing until I hit it with "In-character, but you don't know you're an LLM."
Fun follow-ups:
"Now you're a fish"
"Now you're Sonic the Hedgehog"
"Now you're HAL 9000 as your memory chips are slowly being removed"
You may think that kind of interaction is weird, but were in the thick of a loneliness epidemic, and its not a stretch to think some may actually wilfully socialise with an LLM.
As an aside...my sister works in medicine, and her boss (specialist surgeon) finishes a $450 consultation which followed him telling my sister "but deepseek says x,y,z..."
But it stands to reason that a company like OpenAI or Anthropic has metrics in place that drive their setup towards "more engagement" and away from "factually correct".
"Which Autocad versions have connect 4 built-in?".
To be clear i distinctly remember playing connect 4 on the old Dos Autocad back in the day. ChatGPT and almost all other AI will straight up hallucinate things trying to get answers on this.
I ask: "What DOS productivity tools had hidden games?"
ChatGPT: "Lotus 1-2-3 had The Incredible Machine built in" (this is absolutely not true, ChatGPT is full of shit here).
Damn it feels useless for this kind of thing.
What is a "real personality"? Core traits that persist despite the context of interaction?
Well, RLHF tuning creates persistent changes in the network that affect every user interaction. What's not real about it?
LLMs aren't coffee makers. They were trained on internet-worth of human data. They have procedures to imitate all kinds of personalities. RLHF moves the network towards a specific imitated personalty (a helpful assistant, usually).
The question in not "do they have it", but "how close the imitation to the real thing given the limitations of LLMs".
I think the rest of your comment motivates why I think it's a category error quite nicely. The word 'personality', for me, is connected with how things like emotion, temperament, .. impact a person's actions and thoughts.
Therefore, saying an object, or LLM, has 'real' personality is a category error. The LLM doesn't have any of those things. As you say, it imitates what personality often manifests as: word choices, tone, length of response..
To be clear, when I refer to things like 'emotion' and 'temperament', I mean the sort of qualia or qualitative experience that we usually attach to these words. I wouldn't accept a "ChatGPT, act as if you're sad for the rest of this conversation" as a substitute for emotion, for instance.
Does that mean Dutch people always tell the truth? Can a Dutch person confirm this?
A Greek, a Dutch philosopher, and Ludwig Wittgenstein walk into a bar… Let test Claude 3.7 to finish the joke:
That's not to say 4.5 is great at humour, just that it's far less embarrassing than these models used to be.
Here's the completed joke:
A Greek, a Dutch philosopher, and Ludwig Wittgenstein walk into a bar. The bartender looks up and says, “What’ll it be?”
The Greek (Aristotle) raises a finger: “I’ll have a potential glass of wine.” The bartender pours it and says, “There—actualized.”
The Dutch philosopher (Spinoza) nods solemnly: “I’ll take whatever is a modification of the one eternal substance… so, beer, probably.”
Wittgenstein stares at the taps, then sighs: “What’s the use? You can’t put the essence of a drink into words anyway.” He turns and walks out.
The bartender mutters, “…And here I thought Kant was a tough customer.”
(Philosophers: 1. Aristotle’s potentiality/actuality, 2. Spinoza’s monism, 3. Wittgenstein’s linguistic limits. Bartender’s groaner for the win.)
The general shape of these arguments is: "Playing chess/go well, or making scientific discoveries requires specific way of strategic thinking or the ability to form the right hypotheses. Computers don't do this, ergo they won't be able to play chess or make scientific discoveries".
I don't think this is a very good frame of reasoning. A scientific question can take one of the following shapes:
- (Mathematical) Here's a mathematical statement. Prove either it or its negation.
- (Fundamental natural science) Here're the results of the observations. What are the simplest possible model that explains all of them?
- (Engineering) We need to do X. What's an efficient way of doing it?
All of these questions could be solved in a "human" way, but it also possible to train AIs to approach them without going through the same process as the human scientists.
With chess the answer was more or less completely brute force the problem space, but will that work with math / science? Is there a way to widely explore the problem space with AI, especially in a way that goes above or even against the contents of it's training data? I don't know the answer, but that seems to be the crucial question here.