What does Alan Kay think about LLMs?
quora.com
quora.com
"reasoning by correlation" as superstition is a brutal insight.
I don’t think he is right on that one, though. Reasoning by correlation is a kind of empiricism, and it can be tested. “These things happen together” or “if I do this, then that happens” can be disproven with statistical analysis even without any understanding of the underlying mechanisms at play their causes.
Superstition is beyond that; it is a belief that cannot be proven and that is not based on logic. The problem with superstition is not that there is no causal link between the supposed cause and the effect, it’s that there is not even a correlation.
(One may also link fashion/accusations of fashionability etc etc to superstition..)
We do know that many people even in professional fields confuse correlation with causation. And even when that doesn't happen, when only doing measurments, correlation is often considered good enough without giving any thought to the underlying mechanism. This may be the root of Goodhart's law stating that when a measure becomes a target it stops being a good measure. The measure was only correlated with the behaviour the person doing the measuring thought was measuring.
Superstition is similar, but drops any shred of statistical rigour, relying on mere anecdotal evidence.
If you always rub your lucky penny before playing a sports game a superstitious person would account wins for the penny and excuse losses by not having rubbed the penny enough, the right way, not concentrated enough. This also feeds into correlation.
We are not talking about things actually being correlated, we are talking about people believing they are. Those are two wildly different things one is physical reality and doesn't change when you ignore it, the other is bound to individuals and tbeir world view and can take on absolutely ridiculous forms.
I can’t wait for the day when ultra long context LLMs are able to point out omissions and factual contradictions in human produced fake news for any given event
> mjburgess 9 days ago
Suppose you touch a fireplace once, do you touch it again? No.
OK, here's something much stranger. Suppose you see your friend touch the fireplace, he recoils in pain. Do you touch it? No.
Hmm... whence statistics? There is no frequency association here, in either case. And in the second, even no experience of the fireplace.
The entire history of science is supposed to be about the failure of statistics to produce explanations. It is a great sin that we have allowed pseudosciences to flourish in which this lesson isnt even understood; and worse, to allow statistical showmen with their magic lanterns to preach on the scientific method. To a point where it seems, almost, science as an ideal has been completely lost.
The entire point was to throw away entirely our reliance on frequency and association -- this is ancient superstition. And instead, to explain the world by necessary mechanisms born of causal properties which interact in complex ways that can never uniquely reveal themselves by direct measurement.
All I will say is that for people who want to understand his perspective, there's a large epistemological load to overcome. Sampling his talks is a good starting point though: https://tinlizzie.org/IA/index.php/Talks_by_Alan_Kay
My line of thinking is that without an audit trail of thoughts there won't be any trustability. I'm unable to describe in detail what any one specific thing leads to trustability. I can say that being able to demonstrate where answers come from would work for me.
As for LLMs I think that long form typeahead is the right way to think about it.
> The internet scales well. So imagine building entire computing systems this way.
That’s very similar to microservices, and look how messy these tend to get. The internet scales well because it’s a network with flexible capacity and attached processing power, and because its clients are either human users who can deal with failure modes, or software with stable and well-specified point-to-point protocols.
I bet if he simply said "use Erlang", 99% of headlines and discussion would be "Alan Kay said Erlang is the greatest language evah!"
I do appreciate that he respects his audience, ie me, enough that he thinks we can read "message passing is good" and go from there to choosing a message passing language that is suitable for our needs, or even using a non message passing language due to other factors, but recognizing that OOP is about message passing and not state encapsulation, which would impact how I code even in Java.
How is inventing Smalltalk and putting his ideas into reality just "waxing around" or "complaining"?
Oh yeah? Then how do you explain this lecture where he says explicitly that Erlang should be the modern programmer's assembly?
In my experience, people who criticize Kay for being too vague or kvetching too pointlessly or lacking practical experience in computer science today almost always have not bothered to familiarize themselves with the vast majority of his work before deciding their opinion on it. It's hard to disagree with him that this makes us more like a pop culture than a profession after seeing it.
A whole operating system like this would allow you to hack on the user interface and change things on the fly. All of your software would also work like this, promoting open extensibility by the user. You ought to be able to click on any window or other user interface element and be able to view and modify the running code live, without even restarting the computation it’s running (never mind the whole program or even rebooting the computer).
This sort of ability to do live hacking on the internals of what you’re working with is how computers used to work in those early days. It’s also how machinery has always worked in the past. A mechanic could lock in the timing of an engine by rotating the distributor cap and listening to how smoothly it fires.
Our hardware is sufficiently capable enough these days that I'm curious if we can do this to conventional Smalltalk and Common Lisp system/machine designs to re-imagine them, and bring back a level of tight feedback looped developer experience that has gone underground in the mainstream.
If the sender is calling a method, then that is a static arrangement and the result is not different from a function call.
Reminds me of the RESTful guy, Roy Fielding (though I respect and like Alan Kay way more). If you've read any of his online interactions, apparently nobody does REST like he envisioned. It simply doesn't exist in the wild. When asked about some REST API, he'll claim it's not RESTful, it's wrong, that's not what he meant, etc. "But what would be right?", you ask, and he'll reply "Read my paper! HATEOAS! (Hypermedia as the engine of application state)". "Ok, but what does it mean in practice, how do I go about building an API with HATEOAS?"...
...and there's no answer for that besides "read my dissertation". He can only tell you what you're doing wrong, but nobody has been able to build a "true" RESTful API to Roy Fielding's satisfaction.
I think you're confused about who Fielding is and what REST is about. His writing isn't about building "true" RESTful APIs. It isn't really about APIs per se at all, RESTful or not. (Further, Fielding to my knowledge isn't even known for his use of the word "RESTful".) It sounds like you've probably been misled by a bunch of people who aren't Fielding about what's in his dissertation. Not unreasonable in those circumstances that one be told to actually read the thing instead of guessing at the gestalt of it based on ambient chatter which is by and large very misinformed.
Having said that, Fielding's writing is not something I would exalt for its clarity. I'm partial to jcrites's explanatory powers: "REST describes how the Web works" <https://news.ycombinator.com/item?id=23672561>
Hateoas didn't invent anything. He literally just described how web 1.0 worked.
> After 16 years of continuous research and important contributions toward its mission - "Improve 'powerful ideas education' for the world's children and to advance the state of systems research and personal computing" - Viewpoints Research Institute concluded its operations at the beginning of 2018.
The papers and demos are the output. They were a research outfit, not a startup. OP claimed that Kay "never offers anything but the most vague suggestions" when he (et. al.) provided several very concrete working demos. It's hardly his fault if no one took him up on them, is it? (FWIW I suspect some of their work had an effect on MS Word & Excel UI but I don't know for sure.)
> I haven't seen any significant new insights or advances.
Where have you looked? Did you read the papers that VPRI published? OMeta has been mentioned else-thread, I like that Nile programming language, the COLA system seems neat.
Yes. And yet, he created OOP [1]. Strange.
________________
[1] Not on his own.
The sad truth is that most great ideas are doomed to drown in a sea of mediocrity and misinterpretation.
Where it becomes more dangerous, at least in my opinion, is when the LLM only gets it sort of wrong. Maybe the answer it gives you is old, maybe it’s inefficient, maybe it’s insecure or a range of other things, and if you’re new to programming, you’re probably not going to notice. Hell, I’ve reviewed code from senior programmers that pulled in deprecated things with massive security vulnerabilities and never noticed because they were too focused on fast delivery and “it worked”. I can’t imagine how that would work out for people trying to actually learn things.
I’m not sure what we can really do about it though. I work a side gig as an external examiner for CS students. A lot of the curriculum being taught (at least here in Denmark) are things I’ve seen the industry move “beyond” in the previous 20 years. Some of it is so dated that it really makes no sense at all. Which isn’t exactly a great alternative to the LLMs, and it’s only natural that a lot of people simply turn to these powerful tools.
I tend to tell people to ask their favorite LLM to help them solve a crossword. When you ask it to give you words ending on “ing” it’ll give you words that don’t end on “ing” because of how the tokens used work. This tends to be an eye opener for people in regards to how much they trust their LLM. At least until they get refined enough that they can also do these things.
Anyway, it’s a good answer.
Idk, ymmv, but GPT4 beats the pants off of 3.5.
My experience with it for Powershell hasn’t been as impressive as yours. It simply invented functions for the PnP module. Which I suppose isn’t in too much use these days, especially not against on-prem SharePoint. I’m not an expert on Powershell, however, and I really don’t think the documentation on things like PnP search queries is very good. So this was an area where I got to experience what it’s like to use these tools when you’re not the expert. In the end it was trial and error along with various shitty internet articles that helped me.
That being said. You can easily have GPT4 solve your CS exam questions, and likely whatever technical interview you’re given if you’re applying at jobs which do these things. And as such, I guess you could also argue that a lot of what we teach in CS is now even more dated than before. Because even if some of it is sort of useful for basic understanding, I’d bet money on students “cheating” their way through where they can. Because why wouldn’t they use the “calculator”?
Okay, I didn't expect usable results here but everyone should be aware... when even non-mainstream technologies are being discriminated, what about the human spectrum?
I really like this quote. We simultaneously value trust and community, yet so many people also treat it as just another resource to turn into money and power. Alan Kay is a real gem.
What I believe he is getting at is people are going to use LLMs to build systems at scale to further strip mine society.
The "Spaceship Earth" problem is a reference to Limits to Growth. For those who haven't read "Limits to Growth", and the more recent Re-calibration of Limits to Growth, I implore you to do so.
It's a bad thing for sure, but doesn't seem specific to LLMs, and a solution would not be specific either.
When we landed on the moon, people thought we would soon have Mars bases or even colonize other solar systems, yet here we are, 60 years later. It's naive to think every technology goes strictly on an exponential upward curve. Reality isn't that simple or repetitive, really.
And, no, it's not "different this time". Every time people think it is different this time, and it never is. We don't have flying cars, hypersonic jets, molecular assemblers, AGI. We won't have them tomorrow either, just because Sam Altman says so.
As I see it, LLMs have the potential to either reduce the amount of "bullshit jobs/tasks", reduce the amount of time programmers spend on boilerplate, etc.
But those inefficiencies are 95% human made, on purpose. Bullshit tasks are made up because middle managers want more underlings. Languages/frameworks with lots of boilerplate exist because companies would rather hire fifty average programmers rather than ten brilliant ones.
Even if LLMs have the potential to make these things more efficient, the people holding the money bags don't want the inefficiencies removed.
Consider influencers. What they post on social media could 100% be replaced by the outputs of LLMs and diffusion models. It would be vastly more efficient for companies to advertise their products by creating text and images about a pretty person using their products in an exotic location, than to pay for airplane tickets and hotels and salary for an influencer. Yet we don't see even a hint of the "influencer revenue crisis" that this would cause.
Not instantly and not perfectly though. Inefficiency can be very resistant to change as you are pointing out.
1: 100% transparency. Open Source code, fully (and correctly) attributed training data.
2: A predictable model of what these models are actually encoding (so that hypothetical new models (or modifications) can be reasoned about).
For example, the proof of the absence of solution in SAT should be accompanied with the easily verifiable chain of reasoning. This shows the absence of incorrect deductions and missing assignments. Another example is autovectorization in contemporary compilers, they can show you why parts of your loops are not eligible for vectoriztion.
All LM's can do is to show me that these parts of those inputs are important for that output, but nothing else. Thus, they cannot be trusted even for minimally critical tasks.
I did both.
One can teach LLM to write good code now and bad code in future (when everyone lower their guards). And no one can prove who and how made LLM to do that. Also, no one can prove formally the absence of such a plant.
I still find it funny we managed to get image generation working so much better than text.
More than anything it makes me sad. ANY amount of critical thinking would tell you all of this is not true, which isn't to say there's NO USE AT ALL for this technology, it certainly exists and has it's applications, and I also do more or less believe someday we'll create digital intelligence, but at the same time... ChatGPT is not that. DALL-E is not that. These systems are interesting and they have uses but they are not emergent intelligence, they don't know anything, they just assemble words from massive probability matrices and then the people who read those words ascribe meaning to them that is far, far beyond what originated them.
In this way it's not so dissimilar from any garden variety religion, it's just religion for people who think they're too smart to fall into the trap of motivated reasoning and magical thinking.
In a sense, LLMs have reinvented cold-reading from first-principles and created the cleverest Hans of them all.
Or to say, "It's Plato's cave all the way down"
I'm not entirely sure what you mean by "cultish community", from my perspective there are a few distinct communities around LLMs, all focusing on different aspects, all excited about different things.
One common theme across all the groups though is that they used an LLM for the first time and their mind ran wild with the possibilities. That first moment when the LLM does something better than you expected, or even completely unexpected. I think most people understand that their imagination might be overactive in that moment. But it's a rare feeling to be surprised by a new technology (at least for me) these days.
On the other end, we have social media platforms where being a pessimistic curmudgeon ends up getting the likes and shares. And it's just easier to be a pessimistic curmudgeon; the vast majority of ideas never work as well in the real world as they do in your head. I'm just as guilty of this as anyone else. But the real problem is that it puts us into tribes. As someone who is very excited about what LLMs are going to bring to our futures, when I see someone post on Mastodon or HN, or wherever, I become defensive and my monkey brain feels the urge to push back. In particular because I think the criticisms generally voiced are not well reasoned or thought out. Your own post has a tone of dismissal, painting a lot of people, all of whom excited about different things, as a cult who is obsessed with their LLM girlfriend. I would agree that anyone today trying to draw some deeper meaning from the outputs of these systems are probably worthy of dismissal, but I don't think that's the vast vast majority of people who are excited about LLMs. And it makes _me_ sad that the extremists are the ones that get to suck all the oxygen out of the conversation.
We're in the beginning days of this new technology. LLMs are good at doing things traditional software isn't, and bad at doing a lot of things computers are traditionally good at. Natural language answer engines and sex bots might have been some of the first obvious applications of LLMs, but I'm willing to bet there are a lot more undiscovered use cases out there. Simon Willison has some great advice, which is for newcomers to try to break the LLM as quickly as they can, get it to lie to you, or do something wrong. Test its limits. That's part of the process! We're going to need some time to figure it all out and make these systems work well for us. I'm a technologist, and exploring this technology is exciting.
And I bring all that up to say: no, the vast majority of people, I don't think expect the computer to come to life and tell them it loves them. I think the vast majority, in fact, don't know a fucking thing about LLMs beyond maybe the rote copy/pasted code it takes to bring one into existence, or if we're being honest, more likely, the websites to put their credit card information into to gain access to one for their use, and that shift in base assumptions I think explains why they speak so incoherently: they do not understand it in any depth, and thereofre, they might think the computer will come to life, because they don't know much about computers in general and to a layman, what an LLM does can indeed look like a vague imitation of life.
Like, I mean this in the nicest way possible though just by virtue of what I'm going to say, it is going to sound mean, but: tons of the really big pro-AI hype people just, clearly, bluntly, full disclosure, do not know shit about LLM. Quite a large slice of that pie also don't know shit about technology in general, or seemingly, much of anything beyond a business degree? But irrespective of that, to the wider world who aren't in this and don't participate in the groups at hand... those are your representatives, by default. The attention economy has produced them and you have my most sincere sympathies for that.
If you care about veracity then image generation works about as well as text. Frequently you can find details of the image that are just bizarrely wrong, such as hands or food or other basic things. It's the same basic problem: there's no intelligence behind what it's doing, it just regurgitates mostly realistic-seeming pixels that are pretty good at fooling the casual viewer.
Really, it's like those moths with eyespots on them: good at fooling the brain's heuristics but obviously not real.
It's also worth mentioning you can run a heavily customized Stable Diffusion setup at home with fairly modest hardware with satisfactory results if you know what you are doing, but anything you can run at home for LLMs in the same hardware is dog slow and actually kind of terrible.
You can understand it.
That’s why I dread the AI future.
This is how I measure Ed-Tech companies. Do they have an awareness that you cannot replace the connection with other human beings that is an essential part of teaching with "facts" or not? If "yes, they have that awareness" how do they mitigate the problem?
It's hard to compare trust of LLMs to other computing, because many of the things that LLMs get wrong and right were previously intractable. You could ask a search engine, but it's certainly no more trustworthy than an LLM, gameable in its own way. The closest might be a knowledge graph or database, which can formally represent some portion of what an LLM represents.
To be fair the relational systems can and will give "no answer" when an LLM (like a search engine) always gives some answer. Certainly an issue!
But this is all in the realm of coming up with answers in a closed system, hardly the only way LLMs can be used. LLMs can also come up with questions, for instance creating queries for a database. Are these trustworthy? Not entirely, but the closest alternative supportive tool is perhaps some query builder...? I have seen expert humans come up with untrustworthy queries as well... misinterpretation of data is easy and common.
That's just one example of how an LLM can be used. If you use an LLM for something that you can directly compare to a non-LLM system, such as speech recognition or intent parsing, it's clear that the LLM is more trustworthy. It can and does do real error correction! That is, you can get higher quality data out of an LLM than you put in. This is not unheard of in computing, but it is uncommon. Internet networking, which Kay refers to, might be an analog... creating reliable connections on top of unreliable connections.
What we don't have right now is systematic approaches to computing with LLMs.
Except for the hallucination problem, sure. How do you ensure the answer you get from an LLM is not a hallucination? It sure doesn't.
There are plenty of bits of data that can be hard references to datasets or actual physics to assign their probability value. And there are other things that can't.
Intent parsing and speech recognition both frequently return inaccurate responses. It is true that with an LLM it can return something that is sensible but invented, where other systems typically return less inventive wrong answers.
With any of these systems you want to balance the probability of an incorrect interpretation against the impact of the action being taken, and get confirmation based on that. That's pretty normal engineering and UX.
What's kind of funny is almost all the people that complain about accuracy of LLMs would gladly answer the question of "What's the chance for rain today" without giving giving you a 5 minute lecture on what forecasts actually mean.
Without access to real-time weather data, I'd estimate that there's approximately a 50-60% chance of rain in Seattle on March 20th, based on historical averages. However, for the most accurate forecast, I'd recommend checking a reliable weather website or app closer to the date." -- chatgpt
(prompted with the current date, location, an admonition to estimate, and a promise that I understand it will likely be wrong)
It was a conscious decision by corporations to implement the dropping of constraints and search terms when too few results would have been returned. Today the search operators are a joke.
One might find ambiguity in his criticism though, that LLM alone are insufficient... but, that's what Ximm's Law is saying. It's not very interesting to (as I would say, he does here) take on straw, rather than steel.
(A steely defense of LLM is to say that no one is particularly interested in scaling LLM without other improvements, though scaling alone provides improvements; multi-modal, multi-language, long-context, and most of all augmented systems which integrate LLM into systems rather than making them "systems on a chop", are where things look and IMO will be interesting.)
> A key part of their design was to not allow direct sending of commands — only bits could be sent. This means that (other) software inside each physical computer has the responsibility to interpret the bits, and the power to do (or not do) some action
seems, at least on a basic reading, to contradict this famous little argument (or maybe trolling?) he had on HN with Rich Hickey where he seems to be suggesting that one shouldn't just send raw bits, but also a little interpreter along with the data: https://news.ycombinator.com/item?id=11945722
Maybe this is an inevitable consequence of always speaking so abstractly/vaguely, but it also makes it difficult to know what exactly he's suggesting the industry, that he is so routinely critical of, should concretely do next.
But we somehow like the idea of putting everything behind an API and then call it a solved problem.
All they can do is be observed by us and with that expose us to get induced by association, ontologically human hallucinations and deal with its outcome of real consequences.