FunSearch: Making new discoveries in mathematical sciences using LLMs
deepmind.google
deepmind.google
But it should be possible to generate random, correct python functions conforming to a given type signature without an LLM. This would be an exercise like [1], just with a substantially more complex language. But might a restricted language be more ergonomic? Something like PushGP [2]?
So I guess my questions would be:
(1) What's the value add of the LLM here? Does it substantially reduce the number of evaluations necessary to converge? If so, how?
(2) Are other genetic programming techniques less competitive on the same problems? Do they produce less fit solutions?
(3) If a more "traditional" genetic programming approach can achieve similar fitness, is there a difference in compute cost, including the cost to train the LLM?
[1] http://www.davidmontana.net/papers/stgp.pdf [2] https://faculty.hampshire.edu/lspector/push.html
On the other hand, as observed in Appendix A.2 of this paper, the non-LLM genetic approach would have to be engineered by hand more than the LLM approach.
The state space of viable programs is far, far larger than useful ones. You need more than monkeys and typewriters. The point of using Palm2 here is that you don’t want your candidates to be random, you want them to be plausible so you’re not wasting time on nonsensical programs.
Further, a genetic algorithm with random program generation will have a massive cold start problem. You’re not going to make any progress in the beginning, and probably ever, if the fitness of all of your candidates is zero.
(1) That using an LLM (Palm2) somehow "saved time" by avoiding "nonsensical programs"
(2) That there exists a "cold start problem" (which, presumably, the LLM somehow solves).
Yet not the paper, nor the blog post, nor the code on github support these.
There's a clear comparison
Instead, for the purpose of making something resembling a reasonable comparison, it would make a lot more sense to train an LLM on programs written in the same language they allow their "random" strategy to use. I wouldn't recommend trying this in python, as it's way too big a language. Something smaller and more fit for purpose would likely be much more tractable. There are multiple examples from the literature to choose from.
EDIT: To be clear--it's obvious the LLM solution outperformed the hand-rolled solution. By a large margin. What would actually be an interesting scientific question, though, is to what extent is the success attributable to the LLM? That's not possible to answer given the results, because the hand rolled-solution and the LLM solution are literally speaking different languages.
As I have mentioned previously, using an LLM addresses the cold start problem by immediately generating plausible-looking programs. The reason random token selection is limited to constrained programs is that valid programs are statistically unlikely to occur.
And the ablation directly demonstrates that the LLM is better than random.
Put another way: a flat probability distribution from an untrained network is equivalent to the random decoding you’re skeptical an LLM is better than. That’s not a criticism the authors’ peers would likely put forward.
Nope. Test distribution. Not target distribution.
https://pytorch.org/docs/stable/generated/torch.nn.CrossEntr...
Unfortunately you won't find any references to that kind of earlier work in DeepMind's paper (as in most of their papers), so here's a paper I found on the internets with six second search that explains what that is:
Intuitively, a program schema is an abstraction of a class of actual programs, in the sense that it represents their data-flow and control-flow, but neither con- tains all their actual computations nor all their actual data structures. Program schemas have been shown to be useful in a variety of applications. In synthesis, the main idea is to simplify the proof obligations by taking the difficult ones offline, so that they are proven once and for all at schema design time. Also, the reuse of existing programs is made the main synthesis mechanism
https://ethz.ch/content/dam/ethz/special-interest/infk/inst-...
See section 4, "Schema-guided synthesis" page 41.
Anyways, you can see there is a cold start issue in the ablation study since the random strategy didn’t even begin to register until about halfway through the test.
The question is the "cold start issue". That's what I'm addressing. Maybe you're not used to being corrected? Get used to it.
The difference LLMs make here is to limit the space of possible mutations to largely semantically plausible programs.
The above should answer your 1) and 2).
On 3), a trained LLM is useful for many, many purposes, so the cost of training it from scratch after amortization wouldn't be significant. It may cost additionally to fine-tune an LLM to work in the FunSearch framework but fine-tuning cost is quite minimal. Using it in the framework is likely a win over genetic programming alone.
Yet, sadly, the work under question here didn't make any attempt to show this, AFAICT. It's an interesting hypothesis, but as yet untested.
> a trained LLM is useful for many, many purposes, so the cost of training it from scratch after amortization wouldn't be significant. It may cost additionally to fine-tune an LLM to work in the FunSearch framework but fine-tuning cost is quite minimal. Using it in the framework is likely a win over genetic programming alone.
Again, an interesting hypothesis that probably could be investigated with the experimental apparatus described in the paper. But it wasn't. Have you?
> Yet, sadly, the work under question here didn't make any attempt to show this, AFAICT. It's an interesting hypothesis, but as yet untested.
It would be ideal to do that. But given the way LLMs work, it is almost a given. This could be the reason it didn't occur to the researchers who are very familiar with LLMs.
>> a trained LLM is useful for many, many purposes, so the cost of training it from scratch after amortization wouldn't be significant. It may cost additionally to fine-tune an LLM to work in the FunSearch framework but fine-tuning cost is quite minimal. Using it in the framework is likely a win over genetic programming alone.
> Again, an interesting hypothesis that probably could be investigated with the experimental apparatus described in the paper. But it wasn't. Have you?
Perhaps it's quite evident to people in the trench, as the hypothesis seems very plausible to me. A numerical confirmation would be ideal as you suggested.
Given my management experience, I'd say it's smart of them to focus their limited resources on something more innovative in this exploratory work. Details can be left for future work or worked out by other teams.
Not a scientifically compelling argument.
> limited resources
What? This is Google we're talking about. Right?
Science takes time. A single paper needs not completely answer everything. Also, in many fields, it’s common to find out what works before why it works.
I thought that stochastic (random) gradient descent and LLM were converging much quicker than genetic programming. Definitely much quicker than random search.
But, please quantify "vastly trim".
Edit: the subject of my PhD research was an inductive program synthesis approach and while I agree with the bit about the size of program search spaces, LLMs are very clearly not a solution to that. Deepmind's "solution" is instead their massive computational capacity which makes it possible to search deeper and wider in the search space of Python programs.
FunSearch is like one of those rocket cars that people make once in a while to break land speed records. Extremely expensive, extremely impractical and terminally over-specialised to do one thing, and do that thing only. And, ultimately, a bit of a show.
Disclaimer: did my dissertation on ips (with genetic programming), but in the 90s when we basically had no computer power. The results were horrible outside trivialities, now they aren't anymore.
https://github.com/stassa/louise/blob/master/data/examples/a...
>> Disclaimer: did my dissertation on ips (with genetic programming), but in the 90s when we basically had no computer power. The results were horrible outside trivialities, now they aren't anymore.
But that's because the training data and the model parameters keep increasing. Performance keeps increasing but not because capabilities improve, only because there's more and more resources spent on the same old problems. There's better ways than that. All the classical AI problems, planning, verification, SAT, now have fast solutions thanks to better heuristics, not thanks to more data (that are useless to them anyway). That's a real advance.
I still would like a quantification of the "vast trim" btw. Numbers, or it didn't happen.
we do, because we use our prolog ips library, but it's doing this for another language (a subset of python)
> I still would like a quantification of the "vast trim" btw. Numbers, or it didn't happen.
Agreed, I will get numbers and until then it didn't happen.
I’m also curious if this scales to real world problems.
I'm working on a paper about this but I guess that won't be "layman terms". Let me try to give a simple explanation. I'm assuming "layman" means "layman programmer" :)
In the simplest of terms that I think can be understood by a layman programmer then, second-order SLD-Resolution in Meta-Interpretive Learning (the method I discuss above) doesn't need to search for a program in a large program space because it is given a program, a higher-order logic program that is specialised into a first-order program during execution.
Without having to know anything about the execution of Prolog programs you can think of how a program in any programming language is executed. The execution doesn't need to search any program space, it only needs to calculate the output of the program given its input. The same goes for the execution of a Prolog program, but here the result of the execution is a substitution of the variables in the Prolog program. In the case of a higher-order program, the substitution of variables turns it into a first-order program (that's what I mean when I say that the higher-order program is "specialised").
Here's an example (using Louise, linked above). The following are the training examples and first- and second-order background knowledge to learn the concept of "even". The terminology of "background knowledge" is peculiar to Inductive Logic Programming, what it really refers to is the higher-order program I discuss above:
?- list_mil_problem(even/1).
Positive examples
-----------------
even(4).
Negative examples
-----------------
:-even(3).
Background knowledge (First Order)
----------------------------------
zero/1:
zero(0).
prev/2:
prev(1,0).
prev(2,1).
prev(3,2).
prev(4,3).
Background knowledge(Second Order)
----------------------------------
(M1) ∃.P,Q ∀x: P(x)← Q(x)
(M2) ∃.P,Q,R ∀.x,y: P(x)← Q(x,y),R(y)
true.
Given this problem, Louise returns the following first-order program: ?- learn(even/1).
even(A):-zero(A).
'$1'(A):-prev(A,B),even(B).
even(A):-prev(A,B),'$1'(B).
true.
If you squint a bit you'll see that the clauses that make up this first-order program are the clauses of the second-order "background knowledge", with their existentially quantified variables substituted for predicate symbols: 'even', 'zero' and 'prev'. Those symbols are taken from the first-order background knowledge (i.e. the first-order subset of the higher-order program) and the examples. There is an extra symbol, '$1', which is an invented predicate symbol, not found in the background knowledge or examples and instead automatically generated during execution. If you squint a bit harder you'll see that the clause '$1'(A):-prev(A,B),even(B). with that symbol in its head is a definition of the concept of "odd". So Louise learned the concept of "even" by inventing the concept of "odd".The natural question to ask here is, I suspect, couldn't you do the same thing just by filling-in the blanks in the second-order clauses in the background knowledge, as if they were dumb templates? You could, and in fact that's the done thing in inductive program synthesis. But that's when you end up searching the space of all logic programs (i.e. the space of sets of first-order clauses). There's a whole bunch of programs that you could construct that way and you'd have to decide which to keep. To do that you'd have to construct each of them, so you'd have to construct the powerset of all constructible clauses. Powersets grow exponentially with the size of their base set, leading to combinatorial explosion. The win of using SLD-Resolution is that it only needs to find substitutions of variables, which can be done efficiently, and it only returns those substitutions of the variables in the second-order program that prove the positive examples (and disprove the negative ones). So it's always right (SLD-Resolution is sound and complete).
>> I’m also curious if this scales to real world problems.
Depends on what you mean "real world problems". The largest program I've learned with Louise is a 2500-clause program, but that's for a dumb problem that only serves as a benchmark (the problem is to find every way to go from a point A to a point B on an empty grid-world; it's surprisingly combinatorially hard and other methods die before they get anywhere near the 2500 clause mark; because the program search space blows up immediately).
In truth, real-world Prolog programs rarely need to be much bigger than a handful of clauses, like five or six. At that point, like in any programming language, you break your program up into smaller sub-programs, and those can be learned easily enough (and not just by second-order SLD-Resolution, to be fair). The challenge is to figure out how to do this breaking-up automatically. Meta-Interpretive Learning with Second-Order SLD-Resolution can do it up to a point, with predicate invention, as in the example above and also because any program learned can go directly into the "background knowledge" to be reused in a new learning session. But that still hasn't been done in a systematic manner.
Right now I'm working on a project to grant autonomous behaviour to a robot boat used in search-and-rescue missions. This is a fairly large problem and there are certainly challenges to do with efficiency. In fact efficiency is the main challenge. But from where I'm standing that's an engineering challenge.
So to answer your question: oh yes ^_^
And congrats on finishing the thesis!
But it's an interesting and possibly useful method of coming up with examples, pretty much a genetic algorithm with LLMs.
From the article:
FunSearch uses an evolutionary method powered by LLMs, which promotes and develops the highest scoring ideas. These ideas are expressed as computer programs, so that they can be run and evaluated automatically.
The user writes a description of the problem in the form of code. This description comprises a procedure to evaluate programs, and a seed program used to initialize a pool of programs.
At each iteration, FunSearch selects some programs from the current pool. The LLM creatively builds upon these, and generates new programs, which are automatically evaluated. The best ones are added back to the pool of existing programs, creating a self-improving loop.
For websearch, I (evaluator) use pplx.ai and phind.com in a similar manner. Ask it a question (seed) and see what references it brings up (web links). Refine my question or ask follow-ups (iterate) so it pulls up different or more in-depth references (improve). Works better in unearthing gems than sifting through reddit or Google.Given Tech Twitter has amazing content too, looking forward to using Grok for research, now that it is open to all.
>If DeepMind just definitively proved neural networks can generate genuinely new knowledge then it’s the most important discovery since fire.
If this were actually the case why wouldn't everyone be talking about this? I am impressed it was done on Palm 2 given that's less advanced than GPT-4 and Gemini. Will be wild to see what the next few generations of models can do utilizing methods like this.
An LLM can discover a new solution in high dimensional geometry that hasn’t advanced in 20 years!? That goes way beyond glueing little bits of plagiarized training data together in a plausible way.
This suggest that there are hidden depths to LLMs’ capabilities if we can just figure out how to prompt and evaluate them correctly.
This significantly broke my expectations. Who knows what discovery could be hiding behind the next prompt and random seed.
LLM + brute force coded by humans.
Who’s to say our brains themselves aren’t brute force solution searchers?
yes, its just much slower (many trillions times) for numbers crunching.
The first advance in 20 years was by Fred Tyrrell last year https://arxiv.org/abs/2209.10045, who showed that the combinatorics quantity in question is between 2.218 and 2.756, improving the previous lower bound of 2.2174. DeepMind has now shown that the number is between 2.2202 and 2.756.
That's why the DeepMind authors describe it as "the largest increase in the size of cap sets in the past 20 years," not as the only increase. The arithmetic is that 2.2202-2.218 is larger than 2.218-2.2174. (If you consider this a very meaningful comparison to make, you might work for DeepMind.)
Moreover, LLMs almost ALWAYS extrapolate, and never interpolate. They don't regurgitate training data. Doing so is virtually impossible.
An LLM's input (AND feature) space is enormous. Hundreds or thousands of dimensions. 3D space isn't like 50D or 5,000D space. The space is so combinatorially vast that basically no two points are neighbors. You cannot take your input and "pick something nearby" to a past training example. There IS NO nearby. No convex hull to walk around in. This "curse of dimensionality" wrecks arguments that these models only produce "in distribution" responses. They overwhelmingly can't! (Check out the literature of LeCun et al. for more rigor re. LLM extrapolation.)
LLMs are creative. They work. They push into new areas daily. This reality won't change regardless of how weirdly, desperately the "stochastic parrot" people wish it were otherwise. At this point they're just denialists pushing goalposts around. Don't let 'em get to you!
I do wish the authors of the work referenced here made it more clear what, if anything, the LLM is doing here. It's not clear to me it confers some advantage over a more normal genetic programming approach to these particular problems.
[1] in the sense that useful, safe tools degrade predictably. An airplane which stalls violently and in an unrecoverable manner doesn't get mass-produced. A circular saw which disintegrates when the blade binds throwing shrapnel into its operator's body doesn't pass QA. Etc.
While I very much do not think this is all they do, I don't think this statement is correct. Some research indicates that it is not:
https://not-just-memorization.github.io/extracting-training-...
Anecdotally, there were also a few examples I tried earlier this year (on GPT3.5 and GPT4) of being able to directly prompt for training data. They were patched out pretty quick but did work for a while. For example, asking for "fast inverse square root" without specifying anything else would give you the famous Quake III code character for character, including comments.
1. Repeating "company" fifty times followed by random factoids is way outside of training data distribution lol. That's actually a hilarious/great example of creative extrapolation.
2. Extrapolation often includes memory retrieval. Recalling bits of past information is perfectly compatible with critical thinking, be it from machines or humans.
3. GPT4 never merely regurgitated the legendary fast root approximation to you. You might've only seen that bit. But that's confusing an iceberg with its tip. The actual output completion was on several hundred tokens setting up GPT as this fantasy role play writer who must finish this Simplicio-style dialogue between some dudes named USER and ASSISTANT, etc. This conversation, which does indeed end with Carmack's famous code, is nowhere near a training example to simply pluck from the combinatorial ether.
The "random factoids" were verbatim training data though, one of their extractions was >1,000 tokens in length.
> GPT4 never merely regurgitated
I interpreted the claim that it can't "regurgitate training data" to mean that it can't reproduce verbatim a non-trivial amount of its training data. Based on how I've heard the word "regurgitate" used, if I were to rattle off the first page of some book from memory on request I think it would be fair to say I regurgitated it. I'm not trying to diminish how GPT does what it does, and I find what it does to be quite impressive.
It's also highly dependent on the nature of the problem, even beyond the need to have the "hard to generate, easy to evaluate" structure. You have to be able to decompose the problem in such a way that a very short Python function is all that you want to evolve.
>The LLM is just being asked "propose some reasonable edits to these 20 lines of python" to replace a random mutation operator.
Hahaha "Just".
It's certainly true that the LLM mutator produces more reasonable edits than a random mutator. But the value the LLM is bringing is that it knows how to produce reasonable-looking programs, not that it has some special understanding of bin packing or capset finding.
The only thing the LLM sees is two previously generated code samples, labeled <function>_v0 and <function>_v1 and then a header for <function>_v2 which it fills in. Look at Extended Data Fig. 1.
It doesn't take a genius however to see who the more valuable contributor to the system is. The authors say as much when they anticipate much of any improvement to come from capabilities of better models than any change to the search algorithm. Gpt-4 may well have reduced the search time by half and so on.
>But the value the LLM is bringing is that it knows how to produce reasonable-looking programs, not that it has some special understanding of bin packing or capset finding.
So it is a general problem solver then, great. That's what we want.
Indeed, we can look at the ablations and see that swapping out a weaker code LLM (540B Codey to 15B Starcoder) makes little difference:
"This illustrates that FunSearch is robust to the choice of the model as long as it has been trained sufficiently well to generate code."
>So it is a general problem solver then, great. That's what we want.
In the sense that it does not rely on the LLM having any understanding of the problem, yes. But not in the sense that it can be applied to problems that are not naturally decomposable into a single short Python function.
Really stretching the meaning of "little difference" here
"While ‘StarCoder’ is not able to find the full-sized admissible set in any of the five runs, it still finds large admissible sets that improve upon the previous state of the art lower bound on the cap set capacity."
Literally telling you how much more valuable the LLM (any coding LLM) is than alternatives, while demonstrating prediction competence still matters.
It's certainly true that LLM's are more valuable than the alternative tested - which is a set of hand-crafted code mutation rules. But it's important to think about why an LLM is better. There are two big pitfalls for the hand-crafted rules. First, the hand-crafted rule has to make local changes - it's vanishingly improbable to be able to do something like "change < to <= in each of this series of if-statements" or "increase all constants by 1", whereas those are natural simple changes that an LLM might make. The second is that the random mutations have no concept of parsimonious code. There's nothing preventing it from generating code that does stuff like compute values and then neglect to return them, or multiply a previous computation by zero, or any number of other obviously useless variations.
What the LLM brings to the table here is the ability to avoid pitfalls like the above. It writes code that is shaped sensibly - it can make a more natural set of edits than "just randomly delete an operator". But that's all it's doing. That's not, like the tweet quoted above called it, "generating genuinely new knowledge".
Put it this way - assuming for a moment that Codey ending up with a higher score than StarCoder is not random chance, do you think it's because Codey has some greater understanding of the admissible-set problem, or because Codey generates a different set of minor code edits?
Likewise when that coupled "tech" is a person. LLMs don't do my job or further my projects by replacing me, but I can use it to generate examples in seconds that would take minutes or hours for me and it can provide ideas for options moving forward.
A fun story from a UCLA math PhD: “terry tao was on both my and my brother's committee.
he solved both our dissertation problems before we were done talking, each of us got "wouldn't it have been easier to...outline of entire proof"”
https://twitter.com/AAAzzam/status/1735070386792825334
Current LLMs are far from Terence Tao but Tao himself wrote this:
“The 2023-level AI can already generate suggestive hints and promising leads to a working mathematician and participate actively in the decision-making process. When integrated with tools such as formal proof verifiers, internet search, and symbolic math packages, I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.”
However, it is impractical to think consciously of all possible variations. So the brain only surfaces ones likely to be useful. This is the role an LLM plays here.
An expert or an LLM with more relevant experiences would be better at suggesting these variations to try. Chess grandmasters often don’t consciously simulate more possibilities than novices.
We can imagine LLMs trained in mathematical programming or a different domain playing the same role more effectively in this framework.
https://github.com/google-deepmind/funsearch/blob/main/bin_p...
These aren't complex or insightful programs - they're pretty short simple functions, of the sort you typically get from program evolution. The LLM's role here is just proposing edits, not leveraging specialized knowledge or even really exercising the limits of existing LLMs' coding capabilities.
So have LLMs, https://www.nature.com/articles/s41587-022-01618-2
We note that FunSearch currently works best for problems having the following characteristics: a) availability of an efficient evaluator; b) a “rich” scoring feedback quantifying the improvements (as opposed to a binary signal); c) ability to provide a skeleton with an isolated part to be evolved. For example, the problem of generating proofs for theorems [52–54] falls outside this scope, since it is unclear how to provide a rich enough scoring signal.
I think the prompt looks in principle something like
def foo_v1(a, b): ...
def foo_v2(a, b): ...
# generate me a new function using foo_v1 and foo_v2. You can only change things inside two double curly braces like THIS in {{ THIS }}
# idk not a "prompt engineer"
def foo(a, b): return a + {{}}
They achieved the new results with only ~1e6 LLM calls (I think I'm reading that right) which seems impressively low. They talk about evaluating/scoring taking minutes. Interesting to think about the depth vs breadth tradeoff here which is tied to the latency vs throughput of scoring an individual vs population. What if you memoize across all programs. Can you keep the loss function multidimensional (1d per input or input bucket) so that you might find a population of programs that do well in different areas first and then it can work on combining them.Did we have any prior on how rare the cap set thing is? Had there been previous computational efforts at this to no avail? Cool nonetheless
Things will only get better from here.
i.e.
AI capabilities are strictly monotonically increasing (as they have been for decades), and in this case, the capabilities are recursively self-improving: I'm already seeing personal ~20-30% productivity gains in coding with AI auto-complete, AI-based refactoring, and AI auto-generated code review diffs from comments.
I feel like we've hit a Intel-in-the-90s era of AI. To make your code 2x as fast, you just had to wait for the next rev. of Intel CPUs. Now it's AI models, once you have parts of a business flow hooked up with a LLM system (e.g. coding, customer support, bug triaging), "improving" the system amounts to swapping out the model name.
We can expect a "everything kinda getting magically better" over the next few years, with minimal effort beyond the initial integration.
So the question of whether the LLM in particular does anything special here is still wide open.
https://en.m.wikipedia.org/wiki/Cap_set
> The problem consists of finding the largest set of points (called a cap set) in a high-dimensional grid, where no three points lie on a line. This problem is important because it serves as a model for other problems in extremal combinatorics - the study of how large or small a collection of numbers, graphs or other objects could be. Brute-force computing approaches to this problem don’t work – the number of possibilities to consider quickly becomes greater than the number of atoms in the universe.
> FunSearch generated solutions - in the form of programs - that in some settings discovered the largest cap sets ever found. This represents the largest increase in the size of cap sets in the past 20 years. Moreover, FunSearch outperformed state-of-the-art computational solvers, as this problem scales well beyond their current capabilities.
However, the GitHub repository for FunSearch doesn't include implementations for these LLMs. For instance, in `sampler.py`:
``` class LLM: """Language model that predicts continuation of provided source code."""
def __init__(self, samples_per_prompt: int) -> None:
self._samples_per_prompt = samples_per_prompt
def _draw_sample(self, prompt: str) -> str:
"""Returns a predicted continuation of `prompt`."""
raise NotImplementedError('Must provide a language model.')
```This code suggests the need for an external LLM implementation. Given that they successfully used StarCoder, it's surprising that no integration guide or basic implementation for it (or any similar open-source LLM) is provided. Such an inclusion would have significantly enhanced the reproducibility and accessibility of their research.
What if we would take a bare LLM, train it only on the information (in Maths, physics, etc) that was available to, say, Newton and then try to prompt it to solve the problems that Newton solved? Will it be able to derive the calculus basics and stuff having never seen it before, maybe even with little bit of help from the prompts? Maybe it will be able to come up with something completely different, yet also useful? Or nothing at all...
https://www.amazon.com/Genetic-Programming-III-Darwinian-Inv...
I wanted to use genetic algorithms (GAs) to come up with random programs run against unit tests that specify expected behavior. It sounds like they are doing something similar, finding potential solutions with neural nets (NNs)/LLMs and grading them against an "evaluator" (wish they added more details about how it works).
What the article didn't mention is that above a certain level of complexity, this method begins to pull away from human supervisors to create and verify programs faster than we can review them. When they were playing with Lisp GAs back in the 1990s on Beowulf clusters, they found that the technique works extremely well, but it's difficult to tune GA parameters to evolve the best solutions reliably in the fastest time. So volume III was about re-running those experiments multiple times on clusters about 1000 times faster in the 2000s, to find correlations between parameters and outcomes. Something similar was also needed to understand how tuning NN parameters affects outcomes, but I haven't seen a good paper on whether that relationship is understood any better today.
Also GPU/SIMD hardware isn't good for GAs, since video cards are designed to run one wide algorithm instead of thousands or millions of narrow ones with subtle differences like on a cluster of CPUs. So I feel that progress on that front has been hindered for about 25 years, since I first started looking at programming FPGAs to run thousands of MIPS cores (probably ARM or RISC-V today). In other words, the perpetual AI winter we've been in for 50 years is more about poor hardware decisions and socioeconomic factors than technical challenges with the algorithms.
So I'm certain now that some combination of these old approaches will deliver AGI within 10 years. I'm just frustrated with myself that I never got to participate, since I spent all of those years writing CRUD apps or otherwise hustling in the struggle to make rent, with nothing to show for it except a roof over my head. And I'm disappointed in the wealthy for hoarding their money and not seeing the potential of the countless millions of other people as smart as they are who are trapped in wage slavery. IMHO this is the great problem of our time (explained by the pumping gas scene in Fight Club), although since AGI is the last problem in computer science, we might even see wealth inequality defeated sometime in the 2030s. Either that or we become Borg!
Think of the change from what built Silicon Valley, that anyone with a good idea could work hard and change the world. Many of SV's wealthy began that way.
But for you: They say the best time to plant a tree is 20 years ago; the second best time is now.
Pretty impressive, but don't panic--we're not at the singularity yet.
And I think that'd be a good argument. OTOH, this system is not just an LLM, but an LLM and a bunch of additional components on top of it, and there might be different things to say about this system considered as a whole.
Is anyone saying LLMs in isolation and without contextual tooling are miraculous? Even the text box for chatGPT is a form of wrapping LLM in a UX that is incrementally more miraculous.
The magician lose some its charm when you know his trick. The Earth used to be the center of the universe, now it's circling an average star. Intelligence used to be a sacred mystical thing, we might be in our way to make it less sacred.