They seem to believe it is a database system. That's really not how this works, and the fact it behaves like one sometimes is disguising what it is doing.
If I say "write a "for" loop from 0 to 10 in Python" probably 50% of implementations by Python programmers will look exactly the same. Some will be retrieving that from memory, but many will be using a generative process that generates the same code, because they've seen and done similar things thousands of times before.
A neural network is doing a similar thing. "Write quicksort" makes it start generating tokens, and the loss function has optimised it to generate them in an order it has seen before.
It's probably seen a decent number of variations of quicksort, so you might get a mix of what it has seen before. For other pieces of code it has only seen one implementation, so it will generate something very similar. There could be local variations (eg, it sees lots of loops, so it might use a different variation) but in general it will be very similar.
But this isn't a database lookup function - it's generative against a loss function.
This is subtle distinction, but it is reasonable that people on HN understand this.
How are these not both the exact same process of memory recollection? Can you elaborate on the difference between memory recall vs a generative process based on conditioning? I understand how these two are different in application, but not understand why one would say they are fundamentally different processes.
The best I can come up with is this:
Imagine you are implementing a system to give the correct answer to the addition of any two numbers between 1 and 100.
One way to implement it would be to build a large database, loaded with "x" and "y" and their sum. Then when you want to find out what 1 + 2 is you do a lookup.
The other method is to implement a "sum" function.
Both give the same results. The first process is a database lookup, the second is akin to a generative process because it's doing calculation to come up with correct result.
This analogy breaks down because a NN does have a token lookup as well. But the probabilistic computation is the major part of how a NN works, not the lookup part.