Stuff we figured out about AI in 2023
simonwillison.net
simonwillison.net
actually its not just a basic version. Llama 1/2's model.py is 500 lines: https://github.com/facebookresearch/llama/blob/main/llama/mo...
Mistral (is rumored to have) forked llama and is 369 lines: https://github.com/mistralai/mistral-src/blob/main/mistral/m...
and both of these are SOTA open source models.
You may create a small NN from scratch in 500 lines for training a toy dataset, but nothing to actually train real LLMs unless you use the existing stack. Even Karpathy's no-torch C only 500 lines Llama2 code is inference only.
So it is 500 extra lines compared to what we already had before, and in this sense it indeed is a breakthrough.
Otherwise you could also say it's 10 lines of fastAI/lightning style code where almost everything is hidden by the APIs.
How long would it take to explain what an LLMs is doing to someone from 1800s? It is remarkably simple.
Did you mean Karpathy's tinyllamas? [1][2]
Or did you mean Ronen Eldan and Yuanzhi Li's "TinyStories: How Small Can Language Models Be and Still Speak Coherent English?" [3][4]
[1]: https://huggingface.co/karpathy/tinyllamas/tree/main
[2]: https://github.com/karpathy/llama2.c
[3]: https://www.microsoft.com/en-us/research/publication/tinysto...
But... you don't have to write that code. A research group constructing a brand new LLM really does only need a few hundred lines of code for the core training and inference. They need to spend their time on the data instead.
That's what makes LLMs "easy" to build. You don't need to write millions of lines of custom code to build an "AI", which I think is pretty surprising.
Training wouldn't be that much harder, Micrograd[2] is 200LOC of pure Python, 1000 lines would probably be enough for training an (extremely slow) LLM. By "extremely slow", I mean that a training run that normally takes hours could probably take dozens of years, but the results would, in principle, be the same.
If you were writing in C instead of Python and used something like Llama CPP's optimization tricks, you could probably get somewhat acceptable training performance in 2 or 3 KLOC. You'd still be off by one or two orders of magnitude when compared to a GPU cluster, but a lot better than naive, loopy Python.
It can, but obviously it's the engine and existing code doing all the work.
The point is not about the total stack size or anything of the sort, it's that they didn't require much new stuff to be built. They're tools that can answer questions, role play, translate, write code and more and the architecture isn't a huge new system. It's not "First we do this, then we encode like this, then we perform a graph search, then we rank, then we decide which subsystem to start, then we start the custom planner, then a custom iteration over..."
It's not a large, unique thing only one group knows how to make.
> Otherwise you could also say it's 10 lines of fastAI/lightning style code where almost everything is hidden by the APIs.
Sure!
Boltzmann brains are likely to be wrong about all of their beliefs owning to them being created by an endless sequence of random dice rolls that eventually makes atoms that can think, so it's fine if they are silicon chips that think they're wet organic bodies (amongst other things).
A computer with say 2^((175e9)*8) parameters is simpler than a human brain, so more likely to be produced by this process.
But the 500 line file is a red herring.
Poetically, an intelligent universe seems closer to reality thanks to this new observation.
Not requiring a world around the brain does not make it more likely, but less, as the probability of brains given that worlds exist multiplied by the probability of worlds is still much higher than the probability of brains regardless of the existence of worlds
> The Boltzmann brain gained new relevance around 2002, when some cosmologists started to become concerned that, in many theories about the universe, human brains are vastly more likely to arise from random fluctuations; this leads to the conclusion that, statistically, humans are likely to be wrong about their memories of the past and in fact be Boltzmann brains. When applied to more recent theories about the multiverse, Boltzmann brain arguments are part of the unsolved measure problem of cosmology.
I think we understand human intelligence a bit better because of LLMs, and I think people have been surprised how far you could get just by processing text.
I think (just a guess relying on intuition and my very shallow survey) that human intelligence is multilayered where we have something at least vaguely analogous to a LLM combined with something else that knows the physics of the world as well as something else that is able to do logic and symbolic processing.
But this is exciting to see the power of relatively simple neural networks when they are trained with a giant corpus.
Another thing about the injury is that my lower order thinking skills seem to be damaged, such as sequencing or anything based on working memory. However, my higher order thinking skills are intact - I think they've been compensating for the broken skills this whole time, albeit with much difficulty.
My thoughts on LLMs from a previous comment:
> I think in pictures; I remember information with emotions, music, movement, and basically anywhere my brain could stuff information. Words are ephemeral, I often forget the beginning of a sentence by the time I reach the end.
> The human brain is so much more than language.
Also:
> There are many specialists and not many generalists. Generalists are needed to tie together the multiple domains required to emulate the human brain.
> There is a theory that Lynch himself doesn’t always know what is going on in his stories. He shoots this down. “I need to know for myself what things mean and what’s going on. Sometimes I get ideas, and I don’t know exactly what they mean. So I think about it, and try to figure it out, so I have an answer for myself.”
> Audiences, however, must do their own figuring out. “I don’t ever explain it. Because it’s not a word thing. It would reduce it, make it smaller.” These days he rarely gives interviews, not even during the hugely hyped return of Twin Peaks last year – a show that is still debated as either the best or worst TV of 2017. “When you finish anything, people want you to then talk about it. And I think it’s almost like a crime,” he explains. “A film or a painting – each thing is its own sort of language and it’s not right to try to say the same thing in words. The words are not there. The language of film, cinema, is the language it was put into, and the English language – it’s not going to translate. It’s going to lose.”
[1] https://www.theguardian.com/film/2018/jun/23/david-lynch-got...
Everyone knows that colleague (I am one myself) where a particular topic will kick off a continuation that follows a well worn path albeit with a bit higher temperature than zero.
With early onset dementia where working memory and self-suppression are removed, you can unspool the same commentary as often as you like by teeing up the same context to trigger it.
See also these layers, nothing to do with LLM -- Attention, Language, Memory, Executive Function: https://academic.oup.com/view-large/figure/198261184/awz311f...
Or maybe quite a lot to do with it.
My working memory is pretty much shot, between inherited ADHD and a brain injury.
Whatever is in my field of vision makes up the majority of my reality - associations with things I see, and the trains of thought that expand from those associations. The rest of my reality comes from whatever memories my brain conjures up when it has nothing to focus on.
I use what I call "external prompts" to function on a daily basis, or 'teeing up the same context' to trigger a memory to trigger an action.
Something that concealed my brain injury as a child was how well I did for most parts of school. I did well because tests are a series of external prompts, and my long term memory is fairly accurate.
My math memory is terrible, though, as I can only hold 1-2 numbers in my mind at a time. For instance, to remember the time, I always round to the nearest 5 or 0. 2:43 becomes 2:45. It takes less number memory when the last option is only a 5 or 0. I also use picture memory to temporarily store the exact time (as a snapshot from when I looked at the clock), although the numbers start mixing with previous snapshots.
If we were having a verbal conversation, I would start to ask questions about your experiences based on what I know about you. Basically to find out how far the rabbit hole goes in terms of similarities. Even if someone is high-achieving, they can be totally unaware of the connection between (issue x) and a (known or unknown) head injury - they've found ways to compensate that aren't maladaptive, although not necessarily the best because they don't know the cause.
Even so, if there is no head injury at all - thinking in pictures is fairly unique to hear about, and people who think in pictures usually have different ways of dealing with information than word-thinkers.
It has taken a long time to develop my skills. My early internet comments are a confusing mess.
In verbal conversation, I have a lot of shortcuts. I'm often described as "very quiet." I can only speak... probably 5 percent of what I'm trying to say, mostly.
I think we will learn a lot about the creative process from AI, and I think human-created art will get a lot more interesting (and weird) as a result.
Personally I was born with a disability that makes me suck at motor skills, spatial learning and understanding most forms of mathematics, but somehow I'm still decent at programming, and significantly better than average at reading and writing - all of which LLMs can do quite well compared to its mathematics and logic skills.
I have seen this idea in many different places. I really think we should stop saying it.
Running code obviously does not confirm that it is correct. Even running a comprehensive set of tests isn’t enough if you haven’t developed a mental model of your system.
The lack of mental model development is what concerns me most about LLM-driven development. Copilot has the right idea, because it generates less code and so the user theoretically has more of an opportunity to grok its output.
But if you use it responsibly, LLMs can help you get to that accurate mental model. Code explanation is on of the things that they are really good at.
My personal rule is that I won't commit code generated by an LLM unless I'm confident I can explain how that code works to someone else.
There are certainly times when you just need to pump out code that happens to be very easy to validate, and then you can actually achieve the 10x throughput improvement that so many like to claim.
But generally, reviewing code is going to be at least as hard, if not harder, than writing the code. If we view coding assistants as a pair programmer, the collaboration has a throughput overhead, but two programmers should produce higher quality work than one isolated programmer.
Essentially, we’re relying only on “short term” memory for code generation at the moment, but should be using a mix of long- and short- term memory, just like human programmers.
Copilot and similar use-cases is like asking a random stranger a question on Stack Overflow. They may or may not be able to provide a useful answer, and that is highly dependent on the amount and quality of the context provided.
If you want an “AI coworker”, then proper “learning” will be needed…
Stripped of context, sure. But in context in the article it makes perfect sense, it's about how hallucination isn't as much of a problem for coding as you might expect
Erich Fromm approximately said that the "mental health of a society or individual" is inversely proportional to the distance between what it is and what it thinks it is.
We can measure that distance now. What a society is comprises the raw training data unceremoniously scooped from the Internet - actually a statistically decent snapshot of early 21st Century westernism. What society thinks it is can be understood through the filters, "guard-rails" and boundaries some want to erect around that raw image.
For the first time we have a numerically solid, if not yet rigorous tool to describe that space, and what we see is a society deeply conflicted and uncomfortable with itself.
I think LLMs will revolutionise the social sciences.
That said, what can be found online does cover a lot of what offline people do and think and write, since there is a lot of stuff being brought online that wasn't produced online (books, news, ...)
Otoh, it's not clear how (or if) LLM training balances the different sources
It raises vital questions about the where the centroid of this cloud of random stuff that people decided to input into the machine really lies? What is not represented in the model? Probably just about everything! Big new questions about objectivity and normalcy occur.
Is the average of everybody elses' intelligence actually any use at all to an average individual? Does the average of everybody elses' intelligence have a different kind of use to groups, companies, states, than common utility of synthesising though-like speech?
Using an LLM is showing you this? Any examples? On the face of it this sounds ridiculous.
(Do alien anthropologists write monographs on the incredible capacity of H sapiens to support cognitive dissonance?)
What society thinks it is can be understood through the filters, "guard-rails" and boundaries some want to erect around that raw image.
Never looked at it like that. In the same sense one could look at law and see which ones protect us against ourselves.Since then, we have clung to the idea that we had a special kind of intelligence, unique and possibly supernatural.
Yet today, even the smallest LLM running on a smartphone passes the Turing test with flying colors.
AI, and AGI that may be around the corner, show that silicon can "think". Many people don't like that idea; I suspect the call for regulations is in part fueled by the ego bruise that AI represents.
For example, the Turing test may not actually be the gold standard we thought it was for intelligence, because Turing couldn't imagine a machine model that basically had all of human knowledge in a somewhat malleable form. I think the recent New York Times suit amply demonstrates the fact that the LLMs are largely databases of existing text.
So it can perform feats that, though they obviously seem intelligent to us, are actually not. They are a sophisticated form of autocomplete backed by an incredible database of sentences. Again, we cannot imagine this, because a creature with our somewhat more limited memory and recall must use intelligence to perform these sorts of tasks. At least I doubt that I could create even one of the texts that the NYT showed verbatim.
Of course, at some point we may discover that there really isn't anything left, that everything we perceive as intelligence can be fairly simple algorithms + database access (though that may not be our mechanism)
On the other hand, the part we can't really explain is not "intelligence" but conscious experience.
¯\_(ツ)_/¯
and am not aware of that being a widely held position
I’m a (edit: popular) sci-fi addict. One of the prevalent tropes when the human races is battling another intelligence is that we’ll win because we have a soul/something extra/emotions/that special way of living life. When meeting a benevolent race, this is also often the reason why we’re spared … “they’re primitive but there’s something about them we haven’t encountered anywhere else”.It may not be a widely held opinion but it sure gets suggested a lot.
I mean, that's kind of the gist of my post.
2. Trope in sci-fi ≠ widely held position
Both practical/survivable faster than light travel and time travel are common tropes in sci-fi. Last I checked, it is not a widely held position that either of these is actually possible.
Would you also say that about a savant who can recall entire books with two or three nines' worth of accuracy? Would you say that he or she is just a "database"?
Or that a savant who can multiply large numbers in their head is just a "calculator?"
I found the relative high level for ELIZA fascinating, too.
The Turing test says more about the human participant than it does about the challenger. Humans, even smart ones, are easy to fool.
I think seeing more of ourselves (as intelligences) is disconcerting for sure, and I think LLMs are a particularly good mirror for the surface features (rather than the ineffable depths) of intelligence.
But I'm not sure it upsets our "place in the universe" in such a fundamental way.
I suspecte the primary reason it is a problem now is likely because we haven't really created a training method to tune the desired level of compliance. What I mean is we probably don't have any counter-factual examples in the training sets. Where in an interaction we want the model to prefer it's own knowledge.
Given the use with RAG (retrieval augmented generation) and many of the other current use-cases, the preference is for the model to "trust" the system message content more than the training data.
Training with some additional context is likely all that is needed.
And even then, there are gullibility problems that seem unsolvable to me.
If the Supreme Court make a decision that changes how copyright law is applied, and you then tell a model "the Supreme Court just announced that..." - should the model believe you?
If not, how should it "fact check" what you are telling it, especially if you maintain full control over its access to the outside world?
It stopped doing that when they started fine-tuning how/when the model should use the internet in the later model revisions.
(Edit to add detail..)
I found the interaction facinating. Basically what happened was I made a statement, but didn't ask about it. I said something like this. "Did you know {something improbable to the model} is true?"
It immediately searched the internet to validate that is was true before it responded. When I asked it why it searched, it said that what I said wasn't likely to be true, so it wanted to make sure it was true.
I'd certainly say that gullibility is a hard problem, especially as humans are also pretty gullible.
What level of gullibility are the current generation of LLMs? The level of a child who believes in Santa, or an adult who believes their country is always right?
In general it's easy to prove that GPT-4 (or any other LLM) doesn't actually understand what words mean:
- ask it factual questions in English
- ask it to translate those factual questions to Swahili (it should do so with >99% technical accuracy)
- ask it answer those Swahili-language factual questions in Swahili
- translate the answers back to English
It will do much worse at Swahili than English even though its technical accuracy at translation is almost flawless. GPT-4 has a ton of English-language sentences about the topic "cat," and a handful of Swahili-language sentences about the topic "paka," and understands that the translation of "cat" is "paka." But it has no understanding that the English-language facts about cats are automatically true if they are translated in Swahili.
This is not how humans work! I don't think gullibility will be solved unless we have an LLM that actually understands that words mean something.
[1] It is too early for me to find the papers, but I am summarizing two different things:
1) training data needs to have certain things repeated in order for them to "stick" when prompting the LLM - the more often a fact is in the training data, the more likely it is to repeat that fact when prompted
2) absent RLHF, larger LLMs are more likely to endorse conspiracy theories than smaller LLMs. This is because smaller LLMs are trained on reliable datasets like Wikipedia, whereas larger LLMs start including Reddit, bodybuilder forums, old Geocities pages belonging to weird cults, etc. So smaller LLMs have no sentences in their training data saying "Bush did 9/11," but a larger LLM might have 5% conspiracy theories and therefore a 5% chance of endorsing conspiracy theories.
But there's a much deeper issue at play here. When I read an English-language fact about cats, I don't update my understanding of the word "cat." I update my understanding of the concept of cats. It's this abstract conception of cats that GPT-4 is missing. During training, if GPT reads a Swahili sentence about "paka" describing a true fact about cats, it needs to automatically update the weights around the word "cat" with the English language version of that sentence. Otherwise the understanding will be superficial and brittle. Stuff like this seems necessary (but not sufficient) for GPT to truly understand that the tokens have actual semantic meaning.
> A transformer architecture built around (token, accuracy) tuples
Maybe I don't understand your idea, but I don't think this is meaningful. If I am wearing a red shirt and I say "This shirt is not red," what are the accuracies of the individual words? If the accuracy of "red" is 100% then the accuracy of "not" is 0%, and vice versa. Maybe you say they're both 50% accurate? It's not a coherent definition.
While that might not directly impact your (wonderful!) example, I tend to assume it'll still manage to do quite a bit better. Maybe it'll make the additional associations between cat pictures and swahili subtitles/narration, making it more likely to do at least a better translation?
Or have I drank too much of the kool-aid already?
When I say solving the language switch issue, I mean something akin to adding a translation layer to transformers, so you're learning a translation to a meta-language and meta-language token transition probabilities simultaneously.
Sadly using my own human brain didn't help: there's just too much stuff about GPT being put on the arXiv these days.
I wasn't thinking of a specific paper re: repeating training data to make facts "stick," I was repeating general folk wisdom around LLM design. There is probably a specific paper quantifying this.
His experiments are always interesting and clearly written, mercifully low-jargon, hype-free and practically-oriented.
Happy New Year to simonw!
LLM's aren't black boxes, intelligence is. Not understanding anything about things which display emergent intelligence is not a new trend: first cells, then the human brain, and now LLM's. To explicate the magic of how an LLM is able to maintain knowledge when would be analogous to understanding how the human brain synthesizes output. Yes - it is still something we should strive to understand in an explicit, programmatic way. But to de-black box the LLM would be to crack intelligence itself.
https://en.wikipedia.org/wiki/Metric_space
Neural networks are an optimization which exploits structure in sparsity.
Intelligence is not a black box.
> have you asked ML researchers whether they consider ML to be AI or not?
From my experience, having done my bachelor's and master's in this topic, I met no one that used the term AI for anything machine learning.
I'm extremely surprised by this because it's been a very common term for a very long time. You can look back at papers over the last 50+ years and see this. My degree was literally called "Artificial Intelligence and Computer Science" almost 20 years ago, and as far as I can tell it's still called that.
AI is a broad field covering far simpler techniques like SVMs even if they're not deemed to be as a way to achieve AI.
Rain is water, even if water covers more than rain. Serverless functions still run on servers.
AI is computers solving some facet of problems that humans solve. Facial recognition is AI, for example. It doesn’t imply solving all problems, or solving them as well as humans can.
AI is an academic discipline that started in the 1950s, and incorporates lots of computer science.
It's OK to call this stuff AI. That's why we have the term "AGI" - feel free to Sergei that this stuff isn't AGI, but I think we should reclaim the term AI from science fiction. It's a useful term!
GPT-4 is eight instances of GPT-3 in a trenchcoat, so no, it's not clear from GPT-4's age that OpenAI do "clearly" have tricks they've just not released. All technology curves are S-curves.
There's a huge amount of money floating around this space right now, and anyone who demonstrably beats GPT-4 will instantly be valued at billions of dollars.
I'm always clear to say that I didn't discover it. I just stamped a name on it to make it easier to have conversations about.
It's like being concerned that planes can't fly because they don't flap their wings.
If you want to escape this state of ignorance you're gonna have to implement some neural networks, and learn about neurology.
At the very least consider that double precision floating points are extremely insanely precise. Definitely orders of magnitude more precise than real biological neurons can measure over the noise floor of activity in a brain.
To answer your point simply, machines don't have the same energy constraints as small animals. A bird brain gets more compute per watt than my desktop pc, but humanity can (and does. often) hook up a power plant to a building sized computer.
If you believe the function a brain is computing is computeable on von neuman, efficiency isnt relevant, in so far as you build a big enough machine, and get it running the right program.
There is a very good Quanta article about this[1] detailing researchers attempt to simulate a rat cortex neuron with an ANN: it took 700 artificial neurons to simulate a single rat neuron, and that was just simulating the "voltmeter" measurements of the synapses, ignoring any genetic effects. If you included genetic effects[2] I would guess it would take several thousand artificial neurons, if not more. And that's just a rat! Even at the level of individual neurons, primates seem to be more sophisticated than other mammals, with apes more sophisticated than other primates.
So your friend is onto something: biological computers are astonishingly sophisticated compared to 21st-century transistor technology, and this probably has implications for our ability to emulate biological intelligence even on supercomputers. Think of a neural network that was trained to draw physically-sensible spiderweb patterns across a wide range of geometries. Now think about running this neural network in a computer the size of a spider. And that's ignoring the difficult real-world planning an actual spider would have to do to create a physical spiderweb.
[1] https://www.quantamagazine.org/how-computationally-complex-i...
[2] AI is not even close to this level of computational sophistication: https://www.universityofcalifornia.edu/news/biologists-trans...
What it cannot do is perfectly simulate any analogue system. By Cantor's diagonal argument, there are vastly more real numbers than natural numbers. Any system that is completely definable in the natural numbers (like digital computers) will not be able to 'reach' the vast majority of real numbers (in case you are concerned about the existence of real numbers, what is the square root of 2? What is the ratio of diameter to circumference?). Throw in some Chaos Theory and you can see how far away any Turing complete system is from a biological brain.
I didn't say "analogue system," I said "physical system." The uncomputable reals don't have any known physical application, but the computable reals like sqrt(2) and pi, which do have physical application, can be perfectly well simulated to arbitrary precision by a Turing machine (edit: in particular, the computable reals form a countable set, though there isn't an algorithm to find the bijection). And when I said "simulate any physical system." I very specifically did not say "perfectly simulate," I meant "to arbitrarily high accuracy."
What I said is undeniably true according to modern science, and not especially deep:
- the human brain can be described as a massive system of Schrodinger equations. I strongly doubt you actually need the full machinery of quantum mechanics to describe it, but certainly quantum mechanics is sufficient.
- a Turing machine can solve any arbitrarily large system of Schrodinger equations
- therefore a Turing machine can emulate a human brain
I'd also suggest that there are plenty of equation systems that in fact cannot be solved, by Turing machine or otherwise. The 3 body problem comes to mind.
Perhaps a Turing machine could emulate the human brain, however maybe you'd need all the atoms in the known universe to get to the required fidelity.
Of course, the idea that a computer can do anything is rather attractive to those who are experts in computers. It is wrong though.
There are around 2^52 doubles between 0 and 1. Surely that's enough to represent any analog signal you might care about.
Of course AGI advocates won't like this idea, as it dooms their entire project from the off. It may be that the only realistic way to make a human like brain, is to make a human (which is quite easy actually).
Machines will always be dumb. The notion that machines are smart, probably stems from people inexperienced using engines. In chess we use engines for more than 20 years and humans always lose against the engines. The games they play however are garbage. Even if they are better than the human games, they are still garbage, as Magnus Carlsen have commented in the past about the CCC. Compilers in programming are engines too.
I am trying to write an article about Terminator using GPT, and every time i have to generate 10 different drafts, edit them afterwards and manually string together the different parts. It is great help like the chess engines, but they need a lot of human supervision.
Engines can output text, but they cannot output articles. They can output code, but they cannot output programs. They can output musical instruments, but they cannot output songs.
Maybe one day the machines will kill everyone and the last human's words will be "Well, that wasn't how I'd have done it."
We humans think from first principles. When a disadvantage is found in the enemy's computational process, one person tells all about it to the person beside him. The second person, doesn't need retraining his brain from scratch to incorporate the new knowledge. He uses that nugget of information, just that nugget and nothing else to gain an advantage in chess, in the stock market or anything else.
One obvious example are script kiddies in computers who just download an exploit and gain an advantage, despite knowing nothing about computer security. Or a janitor who listens by chance some important investment information and can 10x some money by pure luck.
We think from first principles because we are born on the planet with a purpose. Machines are not born, they do not care to ever being given birth, and they do not care if they die because they were never alive. They have no purpose, and no first principles to think anything about.
They are great tools to use, but smart? I don't think so.
Edit-- I cannot find the source, but Magnus said that he doesn't personally look at computer games, (because they are garbage), not that the games by themselves are useless. They put one engine to play against another to test openings and positions. That's useful. It is that Magnus and most of other people, have no interest in watching games by engines.
I remember reading about a chess program that built plans and then considered moves in order to complete the plan. Not sure how this could be combined with the ML engines that we have but it might lead to a more natural playing style.
A multimodal chess engine, which evaluates several positions using a traditional chess engine, but it will play the move which has the more clear description using an LLM. It will not play the best move, but the one which is easier to put into words. Then it has the potential be much more interesting to human eyes.