142 karma · joined May 13, 2022
Normal welding is intentional application of heat to partially melt two parts at the seam, so that they "mix" in semi-liquid state and become one part when they solidify. Welding may or may not use a third material (solder) to aid the process.
For starters, we still need the AI (LLMs for now) to be more efficient, i.e. not require a datacenter to train and deploy. Yes, I know there are tiny models you can run on your home pc, but that's comparing a bycicle to a jet.
Second, for an AGI it meaningfully improve itself, it has to be smarter than not just any one person, but the sum total of all people it took to invent it. Until then no single AI can replace our human tech sphere of activity.
As long as there are limits to how smart an AI can get, there are places where humans can contribute economically. If there is ever to be a singularity, it's going to be a slow one, and large human AI vompanies will be part of the process for many decades still.
I wonder if there is some sort of transition between recalling declarative facts (some of which have been shown to be decodable from activations) on one hand and completing the sentence with the most fitting word on the other hand. The dream that "hallucination" can be eliminated requires that the two states be separable, yet it is not evident to me that these "facts" are at all accessible without a sentence to complete.
If you have one normal sentence and one overly verbose, the latter will have more tokens and therefore more weight.
It's missing a tree structure, so there is no ordering of skills (learn easy stuff before hard stuff because there is a learning curve to anything).
They might be good as prints to hang on your wall, but in the current state they're more "achievement lists" rather than anything resembling a tech tree.
My conjecture is that the LLM "knows" some things that it does not put into words. I don't know what it is, but it seems wasteful to drop the entire state on every token. I even suspect that there is something like a "single logic step" of some conclusions from the context. Though I may be committing the fallacy of thinking in symbolic terms of something that is ultimately statistical.
E.g. if you asked an LLM "think about X and then do Y", if the "think X" part is silent, the LLM has a high chance of:
a) just not doing that, or
b) thinking about it but then forgetting, because the capacity of 'RAM' or neuron activations is unknown but probably less than a few tokens.
Actually, has anyone tried to measure how much non-context data (i.e. new data generated from context data) a LLM can keep "in memory" without writing it down?