We Don't Know How to Compute (2011) [video]
youtube.com
youtube.com
Though we have been building and programming computing machines for about 60 years and have learned a great deal about composition and abstraction, we have just begun to scratch the surface.
A mammalian neuron takes about ten milliseconds to respond to a stimulus. A driver can respond to a visual stimulus in a few hundred milliseconds, and decide an action, such as making a turn. So the computational depth of this behavior is only a few tens of steps. We don't know how to make such a machine, and we wouldn't know how to program it.
The human genome -- the information required to build a human from a single, undifferentiated eukariotic cell -- is about 1GB. The instructions to build a mammal are written in very dense code, and the program is extremely flexible. Only small patches to the human genome are required to build a cow or a dog rather than a human. Bigger patches result in a frog or a snake. We don't have any idea how to make a description of such a complex machine that is both dense and flexible.
New design principles and new linguistic support are needed. I will address this issue and show some ideas that can perhaps get us to the next phase of engineering design.
Gerald Sussman Massachusetts Institute of Technology
It so weird to compare the two. Why would you even want that? We have already plenty of brains on earth. Who sais, that once a computer works like a brain.. it doesn't also come with it's downsides?
There are lot's of problems in compute.. making it more brain like, more linguistic isn't going to solve that.
I believe the whole notion of "a computer as an assistant", is doing more damage than good.
Do people still think there will be any humans left alive in a few hundred years? We've gone as far as we could as biological constructs, switching to a more adaptable architecture is just the next step forward. Our bodies are overfitted for living on this one specific planet which is greatly limiting the real goal of life: to expand to new areas and replicate.
I will go on the record and predict there will be some human beings alive in a few hundred years.
I will go further and boldly predict that some of these future people will like living.
I would hope so too, but here we are, racing full steam towards AGI.
One of the primary "criticisms" of pre-LLM machine learning was that it is inefficient compared to, say, a human, because it requires a plethora of examples to classify something just to match the performance of a human who has only seen a few examples.
That argument no longer holds the same value since the discovery of few-shot learning features of LLMs (though they still have a long way to go). Which makes me think: maybe the reason behind human brain's apparent "efficiency" is the complexity that hides beneath the billions of neurons we have.
Disclaimer: I am an amateur and probably have only a litte idea of what I'm talking about.
A more complex decision such as swerving to avoid a wheel that just came off the car in front, probably takes around ~500 milliseconds, or 50 computational steps.
Considering that modern computers need billions of computational steps to start a calculator app and draw it on the desktop, we have a long way to go.
Though there has been regression on this point, a 1990 NeXTStation Color could do the same with ~50 million cycles on a 1120×832 display. (on 12 MB of RAM and 1.5 MB of VRAM.)
Our visual senses aren't camera feeds - the image we are aware of isn't raw input from the eyes, but rather it's generated out of some combination of raw inputs, low-level hacks to hide limited FOV and discontinuities caused by saccades, and higher-level interpretations based on how we feel, what seems to fit the context based on our prior experiences, etc.
I think it makes sense to assume that this process has many extra off-ramps that can trigger high-level reactions while bypassing conscious awareness. I.e. imagine that half-way through the post-processing stack, the relevant parts of the brain are somewhat convinced you're looking at a fast-moving object about to hit you in the face. There's a point at which it makes sense for the brain to make your body start a dodging move, even if it may be a misprediction, rather than waste a few dozen milliseconds to make sure, at the risk of getting a ball (or a rock, or a fist) up your nose.
Yes, one trick is that. Our computers have to serialize everything. We don't know how to build or program massively parallel computers with distributed memory. But the GP already touched on that point.
Human brain has to do a lot of work processing and integrating all the senses we use, which could cover at least for part of the neural capacity you mention, but I feel that where thinking is concerned, we may have accidentally solved the core problems. I somewhat expect that we'll look back and see that current LLMs were to human minds what the ol' Model T is to Tesla Model S - vastly different in levels of advancement, but also fundamentally the same thing.
Another excellent point. The idea that we may stumble onto massive computational leaps based on imprecise data processing requirements is terrifying. I do think we are nowhere near where we will need to be with respect to computation and, more important, information throughput. But the fact that we "oops"-ed into such a strong statistical model of conversation is deeply unsettling. I do believe GPT based language models will be part of an AGI.
Now, it would be one hell of a coincidence if the model we "oops-ed" into with LLMs, and the model evolution "oops-ed" into, were completely unrelated approaches to the same problem.
For our own brains, this is because evolution is the OG incremental learner - it just doesn't do leaps of faith, nor does it invest resources. Incremental evolutionary advancements must quickly stack into something improving survival rates, or else they get selected out of the gene pool. This strongly suggests that brains (human and animal alike) are based on a design that scaled easily and continuously - one that started simple and could be improved step by step, with each step being simple and conferring some non-zero survival advantage.
For our AI work, LLMs are exactly this: they derive from simple models that were iterated on over the past few decades, in small steps. They are still structurally simple as programs - but they've hit the right structure that allowed them to make qualitative jumps in capabilities from simple scaling - more parameters, more memory, more compute, more training time, more input data. This may be just my own perception, but I feel that LLMs are exactly the kind of model that an evolutionary process would be able to develop incrementally.
With the above in mind, plus the fact that LLMs are trained on the output of our brains (and their output is rated by our brains), makes me believe it's highly likely that, with language models, we've stumbled on the very same fundamental method of implementing thinking that is core to how our own brains work.
IMHO, that makes no sense. Assuming we have a really large state machine with all many possible situations in a state of driving on a straight highway with a few vehicles with more or less constant relative velocity, determining if we need to make a turn is definitely doable in a few computation steps.
We can’t program our AI “models” now because they are based on chains of probabilities. Therefore they can run in real time. They don’t have any real logic that we specifically programmed, only what the model was able to infer and then encode.
Perhaps there should be a marriage between these two attempts of AI.
Hum... The numerical route for AI necessarily has the same complexity in computation time. Those two have equivalent semantics. If you can get a result with one, you can get the same result with the other, it can just differ on how easy it is to program.
By the way, symbolic AI can learn too, you don't need to model things in code.
We as humans have the intelligence to simply see for fractions of a second and then just know — we can know the answers to math problems and how a person is feeling. The symbolic AI programming was an attempt to codify what symbols we are seeing but they are so diverse and nuanced it became apparent this was an ineffective scientific route.
Besides, keep in mind that the currently useful NNs are stuff that requires dozens of GB of memory to run.
Anyway, training a NN is also an attempt to codify what symbols we see on the world.
How do lane keeping assists work?
I’m biased as a molecular biologist but genetics and systems biology are the most incredible thing. Natural nanotech
When I was a kid in the 80's there were a number of gross/gore/horror trading cards for kids - Garbage Pail Kids being the prime example. One of the sets I had was set in some far flung future, I can barely remember the series but part of the theme were living, biological machines. The only card I can clearly remember featured a flying prisoner transport where the prisoners were seated in the ships cavernous stomach and "if upset during transport could accidentally digest the prisoners." (Honorable mention to HL2 featuring "synthtech" - the drop ship, strider and gun ships were living biological machines.)
This stuck in my head as the idea of designing machines which grow like plants or raised like animals was an incredibly amazing and interesting concept. But as I got older and educated the idea them seemed as far flung as the cards themselves. Though reading you post makes me think, maybe they aren't so far flung after all. Though I still fell that a creature which grows steel or titanium skeletons is still quite firmly in the realm of sci-fi.
I don't get how this is impressive or relevant on its own. Sure, it sounds small in face of the final result : a full-grown human. But this 1GB program is born out of and expressed into the full OS that is the physical world.
I remember when desktop apps were written in desktop frameworks in the days when your PC was a Pentium 90 with 64MB of RAM and they opened quickly, they responded immediately to commands and they felt native. Despite the fact that these apps today do very little of value differently than 30 years ago bu the hardware is 1000s of times more powerful, they seem to make untold "metrics" network connections, they have no snap, they feel like massively incorrect abstractions over the hardware, they create weird user-experience and the people who sell them often don't know how to support them.
I feel we have lost our way as an Industry being seduced by what the spoiled rich companies do and applying it to our own little kingdoms instead of doing what we need to do well. We have created false demons to slay and have ended up with such a mess its actually a little embarrassing.
I guess while we lack any recognised/required industry qualifications, like Law or Medicine, we are just all doing our own things for our own reasons and unreasonably expecting everything to keep getting better. I mean, have you used Azure Devops? Who decided that it was reasonable to make a basic web app into a front-end monster that makes 100s of ajax calls?
Well, no, I don't remember that. I remember them locking all the time (up to the point where moving the mouse made them slower) and taking ages to load.
I do agree that it's absurd that the applications that do the same tasks today still take about as long to load and lock about as much, even though computers are 1000s of times more powerful. But they were never fast.
We just expected and did less (at once) with computers at the time. If I close everything on my computer today and only run one application, it's pretty damned responsive (unless it's MS Teams).
The past decades were dull for me in terms of programming languages, compared to the era of Scheme and Prolog. But the computing industry advanced steadily: multi-core programming, map-reduce clusters, CUDA, etc. And these led to these crazy generative neural networks we see today.
I recall GJS gave several talks on this "We Don't Know How to Compute". I can't find the source and I skimmed this video and couldn't find it, but I remember he talked about how Gecko (or Lizard I'm not sure) could be transplanted with extra arms and still function. It strikes me as how much we don't know about computing, if we view biology and other things as computation.
Now, with the AI hype, these frontiers seem open again to exploration. For example, you can give AI agents goals and tools and let them act on their own, it's still clunky but it works and it's improving fast every single day. It's just that we haven't figure out the patterns, implications and best practices yet. Exciting times!
- correctness
- performance
- automation
(We basically still suck at all three.)
As I understood it, the talk mainly addresses the third point. We write inflexible concretions that require a buttload of work, maintenance and resources to achieve mediocre things.
Sussman is pushing programming and software design into directions that break away with this limitation, often through powerful general principles and techniques. The book "Software Design and Flexibility" explores this and provides a variety of examples and advice.
However I think the take on EWD's contributions to computer science could use a different take. On maintainability and extension of systems it is possible to achieve a system like Sussman's propagators while still maintaining proofs of correctness. It's true that overly specific proofs that don't work on the more general principles are harder to maintain in step with a changing program; however the trade-off is that you don't know if the change you made (to the highly dynamic system) will have the desired effect (or at least, you can't prove it doesn't).
For a small circuit diagram it might be simple enough that there are obviously no errors. However the prevailing approach to software development is to make software so complex there are no obvious errors.
Biological systems seem to get around this by encoding a lot of redundant information, preserving a ton of errors, and it takes millions of years to sort them out.
I think though that in some cases it does make sense to trade off some correctness when prototyping and experimenting with ideas. However, in my limited experience, taking that idea to the extreme can lead to systems that only the original author can understand. Mathematics may be impressionistic but the formalization of mathematics, though tedious and pedantic, isn't!
It wouldn't be until a few years later when systems like Lean are shared to the wider world (and Coq continues to improve) that show how maintaining more general proof-level engineering can be done on a larger scale without sacrificing flexibility.
Sussman is legendary. Loved this talk!
Compilers for the most popular languages are pretty amazing these days. Yet it’s still on the developer to make the most of them. One aspect might be performance:
- structure your problem in such a way (carefully manage dependency relationships, eliminate branches, etc). that the compiler can choose to emit vector instructions
- manage memory access patterns (e.g. prefer columnar and SoA access patterns over cache unfriendly pointer chasing structures)
- understand how to maximally exploit cache hierarchies for your problem (e.g. threading to exploit the other L1 caches rather than relying on slower L3 etc.)
Another aspect might be data security rules, e.g. tagging and labelling with the ability to enforce rules for the developer.But there are literally hundreds of aspects to deal with. Our languages today look at the problem space through a toilet roll tube with the other eye closed.
The future of computation is guaranteed to be rich, vibrant and exciting.
There are already a plethora of languages designed for numerical computing. Just use that if that's what you're doing.
[1]: https://dspace.mit.edu/bitstream/handle/1721.1/44215/MIT-CSA...
> It was indeed "experimental epistemology,"
> GJS