Can LLMs Generate Novel Research Ideas?
arxiv.org
arxiv.org
It's great for bouncing ideas off of and getting feedback on them. And yeah, it might product "novel ideas" by mixing and matching existing ideas, but LLMs will never create truly novel ideas. Not in their current form.
The paper didn't really answer the question sadly: their conclusion was just that humans rate LLM answers as more novel than human ones, but less feasible.
Can you give a historic example of a human creating "truly novel ideas" that is not the product of mixing and matching existing ideas?
But I don't think an LLM could ever come up with something like that.
It's such a banal suggestion it makes me think there could be a tension between the requirement to be creative, and the requirement that the next token be among the most probable.
"Hey make me a cool like marketing thing or whatever" isn't going to work.
> Mass–energy equivalence states that all objects having mass, or massive objects, have a corresponding intrinsic energy, even when they are stationary. In the rest frame of an object, where by definition it is motionless and so has no momentum, the mass and energy are equal or they differ only by a constant factor, the speed of light squared (c²).
Now go check the "History" section of that article; which, by the way, takes up about half of it.
I'm trying to come up with other examples, but as the parent said, it then depends how pedantic we want to be e.g. take Poincaré's spacetime, but he'd still be working from two pre-existing ideas, "space" and "time"; the idea of combining the two is (feels?) quite novel and unexpected.
Going back some more, the notion of "space" (or that of "time") feels more primitive, less explainable in terms of other notions.
Which means that an LLM can probably technically come up with novel ideas, since novel ideas aren't some special category of things, it's just that LLMs are not very good at it.
Well, the rest mass notion in special relativity is just a natural derivation. It appears once you have the "real" ideas in place. And those aren't really "existing math" at all. It was a pure physics idea at root: "the universe's laws don't change if you are in motion" (or alternate framings like "you can't tell if you're moving from inside a moving box").
Well, it turns out that if you try to construct such a theory, you end up needing some different (pre-existing) math to help describe it. But the idea isn't math at all.
Then the question is "Can LLMs propose new well-framed, evocative theories like relativity", and the answer is sort of open. In point of fact human beings, to first approximation, can't do this either!
In the end it becomes moot. A novel rearrangement of existing ideas that does something new or different is creativity.
Can LLMs do that? I think they can to a limited extent, but not as well as humans. Is it something we get with scale or does it require a fundamental architectural innovation? Don't know.
The invention of PCR comes to mind:
> During a symposium held for centenarian Albert Hofmann, Hofmann said Mullis had told him that LSD had "helped him develop the polymerase chain reaction that helps amplify specific DNA sequences".
https://en.wikipedia.org/wiki/August_Kekul%C3%A9#Kekul%C3%A9...
https://www.amazon.com/Sounds-Bell-Jar-Psychotic-Authors/dp/...
where the authors (a psychologist and two literature critics) carefully tease apart the connection between creativity and psychosis which is of course problematic because insanity mostly gets in the way of being creative which leads to much more serious definition of what "being creative" really means than one usually finds. (One thing they point out is that a third-rate artist (Andy Warhol?) can become quite prominent if they are good at marketing their work.)
People who are religious will make a theological argument to the effect that "God gave you the power to create when he created you" or "You can be creative because God is inside you".
Atheists may dismiss these arguments out of hand but it's a mistake to do so because of
https://en.wikipedia.org/wiki/Ontological_argument
in the sense that "God" can be defined as "the reason why there is something instead of nothing" which could have no relation to the image of some old patriarch on a throne. If we are made "in his image" we should consider the image revealed in a microscope that reveals that we are based on cybernetic principles that apply to the individual cell as well the whole organism and how those principles apply to the evolution of language and culture as they do to our genetic endowment.
(Insofar as God can delegate his creative ability to you, can't you further delegate it?)
A vulgar version of this is Rodger Penrose's "I can solve math problems because I am a thetan" where he claims to be exempt from the problems that Godel and Tarski and Turing warned you about but since there is nothing complete or consistent about Rodger Penrose these don’t apply (he can't solve Collatz and neither can a OT VIII!)
Muddy thinkers may reject the existence or relevance of God or not explicitly believe they are "a spirit in the material world" but often think there is something uniquely human about creativity (can other animals be creative?) but I sense that the ghost of the arguments above is behind that thinking.
In this transcript I get Python to create something that was never seen before and will never be seen again
>>> uuid4()
UUID('21205a92-2611-4710-b120-4a94f5ccf2d9')
which is by no means interesting; real creativity involves creating something that is useful and/or expressive using certain resources and subject to some system of constraints. Insofar as some task is repeatable, creativity is involved in the creation of some process or and/or system that makes the task repeatable.As Edison put it “Genius is one percent inspiration and ninety-nine percent perspiration” so it is not so interesting that the LLM can generate novel (yuck I hate that word, "novel" is the first word I delete when I have to squash a long paper title to fit into 80 characters) research ideas, I'll be impressed when it can fill out a grant application that gets funded.
LLMs seem to mostly be limited right now by the fact that they're always losing context on new conversations and their interactions don't rewire their neural nets. Hard to come up with a new idea when your brain gets reset with every conversation.
Haha I love this picture!
We shouldn't expect the stochastic parrot to be able to do this though and it is unfair to the stochastic parrot.
It is like expecting a real parrot to say words it has never heard before.
No one asks that of a real parrot because we don't anthropomorphize a real parrot like we do the LLM.
LLMs (unless used in deterministic mode, which you shouldn't anyway) will eventually generate all ideas for the same reason as 1_000_000 monkeys on typewriters will eventually generate War and Peace. The question is only how soon in practice this will happen.
Well, I am nearly certain, that 1_000_000 LLMs (or rather 1_000_000 streams of generation using a few LLMs) will do better than 1_000_000 monkeys on typewriters.
The same 1_000_000 LLMs could do better than 1_000_000 average humans, but we don't know yet.
I have a hard time believing that if we fed an LLM all prehistoric speech uttered from humans that no matter what, it would never escape the paradigms of those people. If that is the case, than would relying on LLMs just get us stuck in our own paradigms and prevent true "progress"?
How much coaxing would it require to get it there? What would that process be like?
Can we learn from that and get an LLM trained on all scientific material pre Einstein and post to discover new physics stuff from what we learn from that process?
The parts of an LLM that teach it language is not disconnected from the parts that teach it facts. Good luck teaching it language with only pre Einstein data.
How much would have to leak in to guide the LLM to the right conclusions? Is it a matter of quantity or quality? Can we successfully eliminate all leakage?
and the authors are saying that this study highlights some of the open problems in building research agents that can generate novel ideas like the llm was bad at self evaluation it couldnt tell which of its own ideas were good or bad and they also found that the llm generated ideas that were too similar to each other lacking diversity i mean thats not surprising right llms are trained on huge datasets but theyre still just pattern recognition machines they dont really understand the context or the implications of what theyre generating
but heres the thing novelty is hard to judge even for experts i mean how do you even define novelty is it just something that nobody has thought of before or is it something that challenges our current understanding of the world and the authors are proposing a follow up study where they actually have researchers execute these ideas into full projects to see if the novelty and feasibility judgements actually translate into meaningful differences in research outcomes which is a great idea i mean thats the only way we can really know if these llms are useful for accelerating scientific discovery or not
anyway im rambling on now but i just think this is a really interesting area of research and im excited to see where it goes can we really use llms to accelerate scientific discovery and what are the limitations of these models and how can we overcome them etc etc
Interesting comment nevertheless, if not somewhat difficult to parse.
This claim is unfalsifiable given common definitions of "thinking" and "coming up with novel ideas".
Here's a simple program that does not even do LLMs that will trivially enumerate all ideas (broken UTF8 handling omitted for brevity):
for (var numeral = new BigInteger(0); ; numeral++)
WriteLine(UTF8.GetString(numeral.ToByteArray());
For any given idea in English LLMs will certainly get it faster.The sun is just a bunch of hydrogen.
A computer is just a bunch of transistors.
Humans are just a bunch of cells.
An artificial neural network is just a bunch of matrix multiplications.
I personally think LLMs are extremely limited and overhyped. But this form of argument seems incorrect to me since it can be used to argue that LLMs cannot do things we already know they can do.
LLMs are like mentally challenged autistic human beings. The fact that we even built such a thing is a milestone in humanity.
Human’s ability is in guiding the prompt (of AIs, and the mind) on what is worth knowing and which hypotheses are worth testing.
Choosing an action from the countable infinite number of actions is the general framework of free will.
I give it only 50% meaningless.
Of course it's all meaningless because most of the input and output is virtually identical. Vary the input heavily and then you will approach 50%. Of course this is assuming the token vector is truly random in terms of subject matter.
That being said, you could take a pre-2018 dataset and test it for discoveries/insights circa 2019 or 2020 and see how well it performs.
If I can't figure out where the plot could possibly go from here, go back and look for where I sent it off the rails earlier such that there's nowhere to take it now; if I can't find dialogue or action that fits, go back and find where I put the character in a situation they'd never get themselves in, or miswrote them to respond to it in a way they never would. Stuff like that, especially once it's had you stuck too long for inspiration latency still to be tenable as the cause.
when discussing Trinity and Neo's visions, the Oracle states: "We can never see past the choices we don't understand".
So while I think LLMs are great fun and I don’t judge anyone using them, personally wanting to use one for a difficult problem is a flag that I should allow more time to think before going ahead.
Plus honestly it's just genuinely fun! I've resisted moving to one of their paid packages just yet because I think I'd spend all day on it!
Back in high school, I put together a program called DreamPool that would just literally pick out a few nouns from a gigantic dictionary file and bubble them up in a little graphic of a well, and then I would sit quietly and spin those concepts around attempting to connect them together.
LLMs are like a version of this on steroids and the potential as a tool for augmentation is huge.
Our thinking is a lot less error prone than that of the LLM's but we have to study for years to absorb prior art that LLM's mostly receive at birth by osmosis. It won't be able to take on the truly stupid ideas but it can combine large numbers of the somewhat stupid.
Like a chess position with lots of possible moves followed by lots of possible responses.
I have the idea for an induced draft umbrella. Stick a fan at the top under an opening.
Is that idea novel? I haven't seen it anywhere, it's just something I came up with. But it's not entirely novel, I'm just borrowing the concept of a fan, and an umbrella.
I don't feel this is entirely out of scope for what an LLM could describe in words?
If you give it phrases from, say, the last five years' worth of published research papers, and it combines phrases at random and spits out the words, yes, in that there will be some interesting research ideas, and maybe even some that a human would not have come up with.
Unfortunately, they'll be buried in a blizzard of stuff that no human would come up with, because they are totally meaningless combinations. Finding the good ones is the issue - and it's not an issue that LLMs can currently solve.
LLMs have helped me generate 1 novel discovery: that every top 10 programming language has a single creator (https://pldb.io/blog/aSingleCreator.html).
They also helped me generate this map yesterday (https://pldb.io/blog/whereInnovation.html), which is the most comprehensive map of programming language creation and software innovation ever created.
It seems pretty silly to claim that Algol 60 has thirteen creators. Sure, Wikipedia lists that many names in some table, but that is not really a significant statistic, now is it? Java 22 probably has hundreds of people who contributed to it.
Fun exercise though :)
The creators define how many creators there were.
It's all open source and anyone can fix any mistakes by updating one line (for example, here's the creators entry for Python: https://github.com/breck7/pldb/blob/b8ae74253733e4aa0fb57d26...).
We generally don't have any disputes but there's the occasional error and always open to pull requests to fix those.
But you do bring up a good point about contributors/maintainers. Adding that data is definitely on the priority list, but might be a 2025/2026 thing.
Perhaps we can use LLMs to invalidate patents. Want to check a patent? Just download the LLM model from before the issue date, then ask the LLM to produce the work. If it succeeds, you have invalidated the patent because the work was not novel.
Regardless of folks opinions on the actual plan, does this idea imply an interesting question?
In some sense all of the ideas in a book already exist. Before I read the book, they haven’t been put through a process of being interpreted by my eyeballs and ingested into my brain. But, they do already exist.
Do the ideas in an LLM latent space (or whatever) already exist? They haven’t been read or interpreted yet, but the sentient cognition that goes into creating the idea has already happened.
What does it mean for an idea to exist anyway?