Rodney Brooks on limitations of generative AI
techcrunch.com
techcrunch.com
He suggests to limit the scope of the AI problem, add manual overrides in case there are unexpected situations, and he (rightly, in my opinion) predicts that the business case for exponentially scaling LLM models isn't there. With that context, I like his iPod example. Apple probably could have made a 3TB iPod to stick to Moore's law for another few years, but after they reached 160GB of music storage there was no usecase where adding more would deliver more benefits than the added costs.
I don’t think the iPod comparison is a valid one. People only have so much time to listen to music. Past a certain point, no one has enough good music they like to put into a 3TB iPod. However, the more data you feed into an LLM, the smarter it should be in the response. Therefore, the scale of iPod storage and LLM context is on completely different curves.
Given that Microsoft this year is all-in on LLM AI, this is surely coming.
But perhaps it will be a premium, paid feature?
I am happy with my comments there, including "MS could devote resources to making MS Teams more useful, but they don't have to, so they don't."
This is not obvious though.
You need to have accurate, properly reasoned information for better decisions.
/sarc
https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_di...
For any given size of system, there will be a ceiling on what it can learn to classify or predict with precision.
I used "system" rather than "model" there for a reason:
Memory in any form, such as context and RAG or API access to anything that can store and retrieve data affects the maximum - a turing machine can be implemented with a very small model + a loop if there's access to an external memory to act like the tape. But if the "tape" is limited, there will be some limitation on what the total system can precisely classify.
Is it that way? For example if it lacks a certain reasoning capability, then more data may not change that. So far LLMs lack useful ideas of truth, it will easily generate untrue statements. We see lots of hacks how to control that, with unconvincing results.
The point is that the more context an LLM or human has, the better decision it can make in theory. I don’t think you can debate this.
Hallucinations and LLM context scale are more engineering problems.
Pretty sure we have enough data for these fundamental tasks.
It's not even remotely enough data for a statistical language processor.
What is happening in the human learning process from those few thousand examples, to deduce so much more about "the rules of math" per marginal datapoint?
How many young children do you genuinely think would do problems like that without messing up a step before having drilled for quite some time?
I'm sure there are aspects to how we generalise that current LLM training processes does not yet capture, but so much of human learning processes involve repeating very basic stuff over and over again and still regularly making trivial mistakes because we keep tripping over stuff we learned how to do right as children but keep failing to apply it with sufficient precision.
Frankly, making average humans do these kind of things consistently right manually even for small numbers without putting a process of extensive checking and revision around it is an unsolved problem. And convincing an average human apply that kind of tedious process consistently is an unsolved problem.
You're overestimating how many examples "drilled for quite some time" represents. In an entire 12 years of public school, you might only do a few thousand addition problems in total. And yet you'll be quite good at arithmetic by the end. In fact, you'll be surprisingly good at arithmetic after your first hundred!
> I'm sure there are aspects to how we generalise that current LLM training processes does not yet capture, but so much of human learning processes involve repeating very basic stuff over and over again and still regularly making trivial mistakes because we keep tripping over stuff we learned how to do right as children but keep failing to apply it with sufficient precision.
LLMs fail when asked to do "short" addition of long numbers "in their heads." And so do kids!
But most of what "teaching addition" to children means, is getting them to translate addition into a long-addition matrix representation of the problem, so they can then work the "long-addition algorithm" one column at a time, marking off columns as they process them.
Presuming they can do that, the majority of the remaining "irreducible" error rate comes from the copying-numbers-into-the-matrix step! (And that can often be solved by teaching kids the "trick" of inserting commas into long numbers that don't already have them, so that they can visually group and cross-check numbers while copying.)
LLMs can be told to do a Chain-of-Thought of running through the whole long-addition algorithm the same way a human would (essentially, saying the same things that a human would think to themselves while doing the long-addition algorithm)... but for sufficiently-large numbers (50 digits, say) they still won't perform within an order-of-magnitude of a human, because "a bag of rotary-position-encoded input tokens with self-attention, where the digits appear first as a token sequence, and then as individual tokens in sentences describing the steps of the operation" is just plain messier — more polluted with unrelated stuff that makes it less possible to apply rigor to "finding your place" (i.e. learn hard rules as discrete 0-or-1 probabilities) — than an arbitrary-width grid of digits representation is.
People — kids or not — when asked to do long addition, would do it "on paper": using a constant back-and-forth between their Chain-of-Thought and their visual field, with the visual field acting as a spatially-indexed memory of the current processing step, where they expect to be able to "look at" a single column, and "load" two digits into their Chain-of-Thought that are indirected by their current visual attention cursor — with their visual field having enough persistence to get them back to where they were in the problem if they glance away; and yet with the ability to arbitrarily refocus the "cursor" in both relative and absolute senses depending on what the Chain-of-Thought says about the problem. Given an unbounded-length "paper" to work on, such a back-and-forth process can be extended to an unbounded-length processing sequence robustly. (Compare/contrast: a Turing machine's tape head.)
Pure LLMs (seq2seq models) cannot "work on paper."
If you consider what is even theoretically possible to "model" inside a feed-forward NN's weights — it can certainly have the successive embedding vectors act as "machine registers" to track 1. a set of finite-state machines, and 2. a set of internal memory cells (where each cell's values are likely represented by O(N) oppositional activations of vector elements representing each possible value the cell can take on.) These abstractions together are likely what allow LLMs to perform as well as they do on bounded-length arithmetic. (They're not memorizing; they're parsing!)
But given the way feed-forward seq2seq NNs work, they need a separate instance of these trained weights, and their commensurate embedding vector elements, for each digit they're going to be processing. Just like a parallel ALU has a separate bit of silicon dedicated to processing each bit of the input registers, an LLM must have a separate independent probability model for the outcome of applying a given operation to each digit-token "touched" on the same layer. Where any of these may be under-trained; and where, if (current, quadratic) self-attention is involved, the hidden-layer embedding-vector growth caused by training to sum really big numbers, would quickly become untenable. (And would likely be doubly wasted, representing the registers for each learned arithmetic operation separately, rather than collapsing down into any kind of shared "accumulator register" abstraction.)
---
That being said: what if LLMs could "work on paper?" How would that work?
For complete generality — to implement arbitrary algorithms requiring unbounded amounts of memory — they'd very likely need to be able to "look at the paper" an unbounded number of times per token output — which essentially means they'd need to be converted at least partially into RNNs (hopefully post-training.) So let's ignore that case; it's a whole architectural can of worms.
Let's look at a more limited case. Assuming you only want the LLM to be able to implement O(N log N) algorithms (which would be the limit for a feed-forward NN, as each NN layer can do O(N) things in parallel, and there are O(log N) layers) — what's the closest you could get to an LLM "working on paper"?
Maybe something like:
• adding an unbounded-size "secondary vector" (like the secondary vector of a LoRA), that isn't touched in each step by self-attention, and that starts out zeroed,
• with a bounded-size "virtual memory mapping" — a dynamic and windowed position-encoding of a subset of the vector into the Q/K vectors at each step, and a dynamic position-encoding of part of the resulting embedding (Q.KT.V) that maps a subset of the embedding vector back into the secondary vector
• where this position-encoding is "dynamic" in that, during training of each layer, that layer has one set of embedding vectors that it learns as being a "input-vocabulary memory descriptor table", describing the virtual-memory mappings of the secondary vector's state-at-layer-N into the pre-attention vector input at layer N [i.e. a matrix you multiply against the secondary vector, then add the result to the pre-attention vector]; and an equivalent "output-vocabulary memory descriptor table", mapping the post-attention embedding vector to writes of the secondary vector [i.e. a matrix you multiply against the post-attention embedding vector, then add to the secondary vector]
• and where the secondary vector is windowed, in that both memory-descriptor-table matrices are indicating positions in a window — a virtual secondary vector that actually exists as a 1D projection of a conceptually-N-dimensional slice of a physical secondary N-dimensional matrix; where each pre-attention embedding contains 2N elements interpreted as "window bounds" for the N dimensions of the matrix, to derive the secondary vector "virtual memory" from its physical storage matrix; and where each post-attention embedding contains 2N elements either interpreted again as "window bounds" for the next layer; or interpreted as "window commands" to be applied to the window (e.g. specifying arbitrary relative affine transformations of the input matrix, decomposed into separate scaling/translation/rotation elements for each dimension), with the "window bounds" of the next layer then being generated by the host framework by applying the affine transformation to the existing window bounds. (And again, with the output window bounds/windowing command parameters being learned.)
I believe this abstraction would give a feed-forward NN the ability to, once per layer,
1. "focus" on a position on an external "paper";
2. "read" N things from the paper, with each NN node "loading" a weight from a learned position that's effectively relative to the focus position;
3. compute using that info;
4. "write" N things back to new positions relative to the focus position on the paper;
5. "look" at a different focus position for the next layer, relative to the current focus position.
This extension could enable pretty complex internal algorithms. But I dunno, I'm not an ML engineer, I'm just spitballing :)
For example it's common to bury an adversary in paperwork in legal discovery to try to obscure what you don't want them to find.
Humans do not do better with excessive context and it is appearing that although you can go to 2m tokens etc. they don't actually understand. They ONLY do well at "find the needle in the haystack, here's a very specific description of the needle" tasks but nothing that involves simultaneously considering multiple parts of that.
Llama-2, 2T tokens, smarter than a box of rocks
Mistral-7B, 8T tokens, way smarter than llama-2
Llama-3, 15T tokens, smarter than anything a few times its size
Gemma-2, 13T synthetic tokens, slightly better than llama-3
(for the same approximate parameter size)
I think it roughly tracks that moar data = moar betterer.
So, the the usual s-curve, that has an exponential phase, then topping out?
Small models are like arm: you get much of the performance you actually need for common consumer tasks, very cheap to run, and energy efficient.
We need both but I personally spend most of my ML time on small models training and I’m very happy with the results.
Still waiting Microsoft to add a email search to Outlook that isn’t complete garbage. Ideally with a decent UI and presentation of results that isn’t complete garbage.
…why are we hoping that AI will make these products better, when they’re not using conventional methods appropriately, and have been enshittified to shit.
Should I be adding extra random tokens to my prompts to make the LLM "smarter"?
Most of outlook search results aren't even relevant, and it regularly misses things I know are there. Literally the most useless search I've ever had to use.
There is an alternative, Lookeen, which positions itself as LookOut's successor, but I've yet to try it.
https://lookeen.com/solutions/outlook-search/lookout-alterna...
It worked well with thin OSTs too but due to how Outlook and Exchange work, it would have to rebuild the index more often.
Definitely a blast from the past reading the word “Lookeen” but mostly good memories about it. I believe the ADMX integration was pretty decent too.
My take to improve AI output is to heavily curate the data you feed your AI, much the like expert systems of old (which were lauded as "AI" also.) Maybe we can break the vicious circle of "I trained my GPT on billions of Twitter posts and let it write Twitter posts to great sucess", "Hey, me too!"
This is what OpenAI is doing with their relationships with companies like Reddit, News Corp etc:
https://openai.com/index/news-corp-and-openai-sign-landmark-...
Problem is that we have a finite amount of this type of information.
The first category has obviously been pre-filtered to put cheaper resources in simpler problems, as sometimes these projects pays reasonable tech contract rates for 1-2 hours of work to improve only 2-3 conversation turns of a single conversation, and it's clear they usually involve more than one person reviewing the same data.
A lot of money is pouring into that space, and the moats in the form of proprietary training data heavily curated by experts is going to be growing rapidly given how much cash the big players have.
They are already training their models https://slack.com/intl/en-gb/trust/data-management/privacy-p...
> Microsoft to integrate an LLM into outlook
Unlikely to happen. Orgs that use MS products do not want content of emails leaking and LLMs leak. There is a real danger that an LLM will include information in the summary that does not come from the original email thread, but from other emails the model was trained on. You could learn from the summary that you are going to get fired, even though that was not a part of the original conversation. HR doesn't like that.
> However, the more data you feed into an LLM, the smarter it should be in the response
Not necessarily. At some point you are going to run out of current data and you might be tempted to feed it past data, except that data may be of poor quality or simply wrong. Since LLMs cannot tell good data from bad, they happily accept both leading to useless outputs.
Didn't they already do this? A friend of mine showed me his outlook where he could search all emails, docs, and video calls and ask it questions. To be fair, he and I asked it questions about a video call and a doc - but not any emails, we only searched emails.
This was last week amd it worked "mostly OK," but having a q/a conversation with a long email feels inevitable
Right now AFAICT the latter would require the full text of all the emails you've ever sent, to be stuffed into the context window together.
There could be separate personal fine-tunes per user, trained (in the cloud) on the contents of that user's mail database, which therefore have knowledge of exactly the mail that particular user can access, and nothing else.
AFAICT this is essentially what Apple is claiming they're doing to power their own "on-device contextual querying."
Yes, but that contradicts the earlier claim that giving AI more info makes it better. If fact, those who send few emails or just joined may see worse results due to the lack of data. LLMs really make us hard to come up with ideas on how to solve problems that did not exist without LLMs.
That’s my biggest worry. I’ve set up my own RAG and it’s kind of sad how even that is not super accurate.
https://support.microsoft.com/en-us/office/summarize-an-emai...
It's a very weird comparison, as putting more music tracks to your iPod doesn't make them sound better, while giving a LLM more parameters/computing power make it smarter.
Honestly it sounds like a typical "I've drawn my conclusion, and now I only need an analogy that remotely supports my conclusion" way of thinking.
In fact, you would then tend to go the other way — once you get the ML model to "solve the problem", you want to then find the smallest and most efficient such model that solves the problem. I.e. the model that is "as stupid as possible", while still being very good at this one thing.
The pretrained model is where the enterprise gold is at.
But the companies building the models past the tipping point scale for that value to be derived are walling up their pretrained model behind very heavy handed fine tuning that strips away most of the business value.
The engineers themselves seem to lack the imagination for the business cases, and the enterprise market doesn't have access to start discovering the applications outside of 'chatbot,' particularly with large context windows of proprietary data fed into SotA pretrained models.
There's maybe a handful of people who actually realize what value is being left on the table, and I think most of them are smart enough not to currently be in positions to make it happen.
IF X > 600 || X < 1 THEN VX = -VX + RAND
IF Y > 600 || Y < 1 THEN VY = -VY + RAND 1. Go in a random direction in a straight line for a while.
2. Then spiral outwards until hit something.
3. Try turning some random angle. If repeatedly hitting a wall, go into wall following mode for a while. Otherwise, go back to step 1.
If you run this long enough to travel over 2x the actual floor area, the odds of achieving near full coverage are pretty good. Who needs navigation?I think Brooks' opinions will age poorly, but if anyone doesn't already know all the arguments for that they aren't interested in learning them now. This quote seems more interesting to me.
Didn't the iPod switch from HDD to SSD at some point, and they focused on shrinking them rather than upping storage size? I think the quality of iPods have been growing exponentially, we've just seen some tech upgrades on other axises that Apple thinks are more important. AFAIK, looking at Wikipedia, they're discontinuing iPods in favour of iPhones where we can get a 1TB model and the disk size trend is still exponential.
The original criticism of the iPod was that it had less disk space than its competitors and it turned out to be because consumers were paying for other things. Overall I don't think this is a fair argument against exponential growth.
Do we still need to make a fair claim against unrestricted exponential growth in 2024? Exponential growth claims have been made countless times in the past, and never delivered. Studies like the Limits to Growth report (1972) have shown the impossibility of unrestricted growth, and the underlying science can be used in other domains to show the same. There is no question that exponential growth doesn't exist, the only interesting question is how to locate the inflexion point.
Apparently the only limitless resource is gullible people.
But one did though? Pronouncements of the future do not invalidate current results.
The same is true for LLMs. What they can do is impressive but to fix what they can't will be hard or even impossible.
There was of course never a recognized success of FSD, since we don't have FSD, but when people started seeing cars drive themselves on public streets 2010 they assumed we would have FSD in 5-10 years, but we still barely have restricted self driving cars today 14 years later.
99.9% of population never saw a self-driving car in their life, and only a few would treat a single spotting on media as indication of anything.
ChatGPT on the other hand is widespread.
Building cars is just harder than deploying a web service.
160 Tera Bytes to store what?
I probably would never listen to more than a few GBs of high fidelity music in my life time.
Why would anyone keep investing in exponential growth or even linear growth beyond a point of utility and economic sense.
This can be extrapolated to miniaturization also ... why is a personal computer not smaller than my fingertip already?
Maybe we store a local copy of a personal universe for you to test ideas out in, I dunno. There'll be something.
> I probably would never listen to more than a few GBs of high fidelity music in my life time.
Well from that I'd predict that uses would be found other than music. My "music" folder has made it up to 50GB because I've taken to storing a few movies in it. But games can quickly add up to TB of media if the space is available.
Whatever innovative use cases people could come up with to store TBs of data in physical portable format is served by these. And along the way the world has shifted to storing data more cheaply and conveniently on the cloud instead.
IPod being a special purpose device , with premium pricing (for average global consumer) and proprietary connectors and software would not have made a compelling economic case to over-spec it in hopes that some unforseen killer use case might emerge from the market.
Well, yes but if we're literally talking the iPod, it has been discontinued. Because it has been replaced by devices - with larger storage, I might note - that do more. I'm working from the assumption that as far as Brooks was talking iPhones basically are iPods. This is why the argument that the tech capped out seems suspect to me. The tech kept improving exponentially and we're still seeing storage space in the iPod niche doubling every few years.
FWIW, I currently have about 100 GB of music on my phone. And that is in a fairly high quality AAC format. Converted to lossless, it might be about five times that size? I don't even think of my music collection as all that extensive. But still, 160 TB would be a different ball game altogether. For sure, there is no mass market for music players with that sort of capacity. (Especially now that streaming is taking over.)
Or they are too young still, or they just got interested in the subject, or, or, or…
Don’t dismiss someone offhand because they disagree with you, they may really have never heard your argument.
> AFAIK, looking at Wikipedia, they're discontinuing iPods in favour of iPhones where we can get a 1TB model and the disk size trend is still exponential.
iPods were single-purpose while iPhones are general computers. While music file sizes have been fairly consistent for a while, you can keep adding more apps and photos. The former become larger as new features are added (and as companies stop caring about optimisations) while the latter become larger with better cameras and keep growing in number as the person lives.
If you haven’t noticed how growth in storage capacity of HDDs in general screeched to a relative halt around fifteen years ago? The doubling period used to be about 12-14 months; every three or four years the unit cost of capacity would decrease 90%. This continued through the 90s and early 2000s, and then it started slowing down. A lot. In 2005 I bought a 250 GB HDD; at the same price I’d now get something like a 15 EB drive if the time constant had stayed, well, constant.
There is, of course, always a multitude of interdependent variables to optimize, and you can always say that growth of X slowed down because priorities changed to optimize Y instead. But why did the priorities change? Almost certainly at least partislly because further optimization of X was becoming expensive or impractical.
> quality has been growing exponentially
That’s an entirely meaningless and nonsensical statement unless you have some rigorous way to quantify quality.
Most rigorous definitions of quality incorporate a target expectation that is impossible to exceed. You are removing errors, so best case, exponential improvement in quality adds nines - a la the sigmoid function you mention.
We have a qualitative measure of quantity, that was where the discussion started - size in bytes.
Storage size in bytes on a pocket iDevice seems to be growing exponentially. Brooks said it stopped doubling and I don't think that is true. iDevices are still doubling their storage size every 2-4 years or so and have been for a while. There was a one-time change to SSDs where they lost a few generations and they are only up to 1TB instead of 160TB (suggesting SSDs are around 7 generations behind HDD which seems reasonable on the face of it to me). But apart from that it has been a pretty steady doubling every few years.
His claim was that iPods didn't get to 160TB because it "nobody actually needed more than that" and that observation is misunderstanding what happened. Apple switched from HDD to SSD because SSD is a much better fit and put the rate of growth back by a few generations but the growth is still ongoing and probably will reach 100s of TB sooner or later.
He was right that iPods didn't need >a few gig of storage ... but that just meant Apple discontinued the iPod brand and replaced it with iPhones, where there is no usage barrier to consuming terabytes of storage. They were obsoleted by the very trend he claimed was over! It appears Apple thinks people did need more, because they aren't making iPods any more.
If you consider LLMs as the iPod of ML, what would the iPhone equivalent be?
the bubble on this is going to make the .com crash look like peanuts.
For ~50 LOC examples ChatGPT can consistently modify the code to add some new parameter or change some behaviour etc.
For using new external APIs it hallucinates often - that said, all of my changes to my static Hugo websites: new shortcodes, modifying short codes like "change the list of random articles to only include articles that have the same 'type'" works excellent - are done by ChatGPT without problems.
> .com crash look like peanuts.
I think what people get wrong about the .com crash: It wasn't a technology crash but a crash of overvalued companies and the sudden fear of VCs. Internet usage and new applications just grew and grew. There was no internet technology crash - and many successful companies like Amazon and eBay just kept working - the pet.com's of the world died (and sadly my Wiki/Blog/Onthology startup died too)
One instance of this was when I was able to extend features of a partial std::functional port for the AVR platform and was able to achieve my goals by asking ChatGPT to generate rather complex C++ template code that would have taken me several days to figure out since I'm not a C++ programmer. In about two hours and several back-n-forth between the code and ChatGPT's interface, I was able to integrate the modifications (about 50 LOCs) and save me the daunting and frustrating task of rewriting around 2000 LOCs I had written for the espressif platform. This is what I would have done if ChatGPT wasn't around.
In this context, look-good-but-broken examples are not really a problem when you can identify what is wrong and communicate the problems back to GPT. These cases do not bode well when we are asserting the correctness and autonomy of AI systems, but they are not as problematic when one seeks to augment his own abilities.
It's like getting a flying car and saying "meh, on a highway it's not really much faster than the classical car". Or getting a computer and saying that the calculator app is not faster than actual calculator.
A robot powered by some future LLM may not be much better at moving stuff, but it will be able to follow commands such as "I am going on a vacation, pack my suitcase with all I need" without giving a detailed list.
I don’t need to store all my music on my device. I can have it beamed directly to my ears on-demand.
We just scaled in a slightly different way.
But the sentiment analysis, summaries, and object detection seem incredibly capable and like the actual useful features of LLMs and similar tensor models.
>object detection
This comes from a different class of machine learning models unless I am mistaken?
This kind of strawman "limitations of LLMs" is a bit silly. EVERYONE knows it can't do everything a human can, but the boundaries are very unclear. We definitely don't know what the limitations are. Many people looked at computers in the 70s and saw that they could only do math, suitable to be fancy mechanical accountants. But it turns out you can do a lot with math.
If we never got a model better than the current batch then we still would have a tremendous amount of work to do to really understand its full capabilities.
If you come with a defined problem in hand, a problem selected based on the (very reasonable!) premise that computers cannot understand or operate meaningfully on language or general knowledge, then LLMs might not help that much. Robot warehouse pickers don't have a lot of need for LLMs, but that's the kind of industrial use case where the environment is readily modified to make the task feasible, just like warehouses are designed for forklifts.
Its the opposite problem to the perception of computers in the 70s, early computers were seen by some as too alien to be as useful as a person across most tasks, llms are seen by some as too human to not be as useful as a person across most tasks. They are both wrong in surprisingly complex ways.
Tons of people mistake knowledge and eloquence for intelligence, even here on HN, it isn't a strawman at all.
A large amount of people--many of them with money to invest--seem to believe otherwise.
Sorry for not commenting on the article directly.
> ... Copilot ... is pretty bad at guessing what I exactly want to do
This happens a billion times a month in chatGPT rooms. User comes with a task, maybe gives some references and guidance. The model responds. User gives more guidance. And this iterates for a while. The LLM gets tons of interactive sessions, it can learn how to rank the useful answers higher. This creates a data flywheel where people generate experience and LLMs learn and iteratively improve. LLMs have the tendency to make people bring the world to them, they interact with the real world through us.
OoenAI hasn't even released their best model/capabilities to the public yet. Their text to image capabilities they show in their gpt-4o page are mine-blowing. The (unreleased) voice mode is a huge step up in simulating a person and often quite convincing.
I've been thinking, in Star Trek terms, what if it's not Lt. Cdr. Data, but just the Ship's Computer?
In Ian Banks universe AIs also have avatars, but for the sole benefit of humans, they like to interact with avatars more than with a voice in their head.
This is basically like when programs went from being text based to UI based. Future software will have AI, LLMs, as a component/interface.
The "LLM is the program" is not adequate, reliable, scalable.
Anyone who can not acknowledge this lacks the discernment to understand that when they ask a computer program a question and get an answer no human could have given, that is super intelligence.
Aesthetic is complex subject but cannot be generated if task is more precise.
Can someone translate what that means? I'm struggling to read past that line, since I just can't wrap my head around what a "Panasonic Professor" is. Does that word refer to something other than the corporation in this context?
It’s common for corporations or wealthy individuals to donate large amounts of money to fund prestigious academics.
You have some input, which may include labels or not.
If it does have labels, ML is the automatic creation of a function that maps the examples to the labels, and the function may be arbitrarily complex.
When the function is complex enough to model the world that created the examples, it's also capable of modelling any other input; this is why LLMs can translate between languages without needing examples of the specific to-from language pair in the training set.
The data may also be synthetic based on the rules of the world; this is how AlphaZero beats the best Go and chess players without any examples from human play.
The kitchen sinks that we can throw at ML are incredibly powerful these days. So, we can quickly iterate through different neural net architecture configurations, check accuracy on a few standard datasets and report whatever configuration that does better on Archiv. Some might call it 'p-hacking'.
Training data gives the model an idea what "blue" means, and what "cat" means.
It can now generate sensible output about blue cats, despite blue cats not being normal.
For example asking an LLM a question that nobody has even asked it before. The degree to which it does a good job on those questions is called "generalisation".
"Extrapolation" is the word you're looking for, for a superficial understanding of machine learning. "Average" isn't even remotely close.
Were one to slice the corpus callosum,
and burn away the body,
and poke out the eyes..
and pickle the brain,
and everything else besides...
and attach the few remaining neurones
to a few remaining keys..
then out would come a program
like ChatGPT's --
"do not run it yet!"
the little worker ants say
dying on their hills,
out here each day:
it's only version four --
not five!
Queen Altman has told us:
he's been busy in the hive --
next year the AI will be perfect
it will even self-drive!
Thus comment the ants,
each day and each night:
Gemini will save us,
Llama's a delight!
One more gigawat,
One more crypto coin
One soul to upload
One artist to purloin
One stock price to plummet
Oh thank god
-- it wasn't mine!Don't worry about where it sleeps in your tool cabinet so much. Put that fastidiousness to better use by reaching in with your intellect to see what favorable positions you can tickle the LLM into. None to be found you say? In the tool world an LLM is closer to a mirror than a hammer.
FAST, CHEAP AND OUT OF CONTROL: A ROBOT INVASION OF THE SOLAR SYSTEM:
https://en.wikipedia.org/wiki/Subsumption_architecture
Subsumption architecture is a reactive robotic architecture heavily associated with behavior-based robotics which was very popular in the 1980s and 90s. The term was introduced by Rodney Brooks and colleagues in 1986.[1][2][3] Subsumption has been widely influential in autonomous robotics and elsewhere in real-time AI.
https://en.wikipedia.org/wiki/IRobot
iRobot Corporation is an American technology company that designs and builds consumer robots. It was founded in 1990 by three members of MIT's Artificial Intelligence Lab, who designed robots for space exploration and military defense.[2] The company's products include a range of autonomous home vacuum cleaners (Roomba), floor moppers (Braava), and other autonomous cleaning devices.[3]
Edit: the direct sensory-action coupling idea makes sense from a control perspective (fast interaction loops can compensate for chaotic dynamics in the environment), but we know these days that brains don't work that way, for instance. I wonder how that perspective has changed in robotics since the 90s, do you know?
https://dl.acm.org/doi/10.1145/1056602.1056608
https://donhopkins.com/home/TurtlesAndDefense.pdf
>TURTLES AND DEFENSE
>Introduction
>At Terrapin, we feel that our two main products, the Terrapin Turtle ®, and the Terrapin Logo Language for the Apple II, bring together the fields of robotics and AI to provide hours of entertainment for the whole family. We are sure that an enlightened application of our products can uniquely impact the electronic battlefield of the future. [...]
>Guidance
>The Terrapin Turtle ®, like many missile systems in use today, is wire-guided. It has the wire-guided missile's robustness with respect to ECM, and, unlike beam-riding missiles, or most active-homing systems, it has no radar signature to invite enemy missiles to home in on it or its launch platform. However, the Turtle does not suffer from that bugaboo of wire-guided missiles, i.e., the lack of a fire-and-forget capability.
>Often ground troops are reluctant to use wire-guided antitank weapons because of the need for line-of-sight contact with the target until interception is accomplished. The Turtle requires no such human guidance; once the computer controlling it has been programmed, the Turtle performs its mission without the need of human intervention. Ground troops are left free to scramble for cover. [...]
>Because the Terrapin Turtle ® is computer-controlled, military data processing technicians can write arbitrarily baroque programs that will cause it to do pretty much unpredictable things. Even if an enemy had access to the programs that guided a Turtle Task Team ® , it is quite likely that they would find them impossible to understand, especially if they were written in ADA. In addition, with judicious use of the Turtle's touch sensors, one could, theoretically, program a large group of turtles to simulate Brownian motion. The enemy would hardly attempt to predict the paths of some 10,000 turtles bumping into each other more or less randomly on their way to performing their mission. Furthermore, we believe that the spectacle would have a demoralizing effect on enemy ground troops. [...]
>Munitions
>The Terrapin Turtle ® does not currently incorporate any munitions, but even civilian versions have a downward-defense capability. The Turtle can be programmed to attempt to run over enemy forces on recognizing them, and by raising and lowering its pen at about 10 cycles per second, puncture them to death.
>Turtles can easily be programmed to push objects in a preferred direction. Given this capability, one can easily envision a Turtle discreetly nudging a hand grenade into an enemy camp, and then accelerating quickly away. With the development of ever smaller fission devices, it does not seem unlikely that the Turtle could be used for delivery of tactical nuclear weapons. [...]
It's a fun/interesting device in science fiction, just like the concept of golems (animated beings) are in folk tales. But it's complete nonsense to talk about it as a possibility in the real world so yes, the label of 'machine learning' is a far, far better label to use for this powerful and interesting domain.
> The creation of consciousness is impossible.
That's where I'd start my argument.
Machines can 'learn', given iterative training and some form of memory but they can not think nor understand. That requires consciousness, and the idea that consciousness can be emergent (which it is my understanding that the 'AI' argument rests upon), has never been shown. It is an unproven fantasy.
In my view, and in the view of major religions (Hinduism, Buddhism, etc) plus various philosophers, consciousness is eternal and the only real thing in the universe. All else is illusion.
You don't have to accept that view but you do have to prove that consciousness can be created. An 'existence proof' is not sufficient because existence does not necessarily imply creation.
1. People have different kinds of conscious experience (just talk to other humans to get the picture).
2. Consciousness varies, and can be present or not-present at any given moment (sleep, death, hallucinogenic drugs, anaesthesia).
3. Many things don't have the properties of consciousness that I attribute to my subjective experience (rocks, maybe lifeforms that don't have nerve cells, lots of unknowns here).
Given this, it's obvious that consciousness can be created from non-consciousness, you need merely to have sex and wait 9 months. Add to that the fact that humans weren't a thing a million years ago, for instance, and you have to conclude that it's possible for an optimisation system to produce consciousness eventually (natural selection).
I think pragmatic thinking, not metaphysics, is what will ultimately lead to progress in AI. You haven't engaged with the actual content of my arguments at all - from that perspective, who's the one talking about fantasies really?
Edit: in case I give the mistaken impression that I'm angry - I'm not, thank you for your time. I find it very useful to talk to people with radically different world views to me.
Quite obviously not true (and also untrue in a documented sense in that many great scientists have not been materialists). Science is a method, not a philosophical view of the world.
> it comes from interacting with the world as it is
That the world is materialistic is unproven (and not possible to prove).
You are simply propounding a materialist philosophy here. As I've said before, it's fine to have that philosophy. It's not fine to dogmatically push that view as 'reality' and then extend on that to build and push fantasies about 'artificial intelligence'. Again, we avoid all of this philosophical debate if we simply stick to the term 'machine learning'.
Particularly with Hinduism, its embrace of various philosophies is very broad and includes strains of Materialism, which appears to be the other person's viewpoint, so again, I should have been more careful with my wording.
I would define thinking as the ability to evaluate propositions and evidence and arrive at some potentially useful plan or hypothesis.
I see no reason why a machine could not do this, but without any awareness ("consciousness") of itself doing it.
I also see no fundamental obstacle to true machine awareness either, but given your other responses, we can just disagree on that.
Thinking, understanding, self-aware, awake, integrated sensory stream, alive, emotional, visual, sensory, autonomy, adaptiveness, etc.
And machines can't learn.
AI is what we've called things far simpler for a far longer time. It's what random users know it as. It's what the marketers sell it as. It's what the academics have used for decades.
You can try and make all of them change the term they use, or just understand it's meaning.
If you do succeed, then brace for a new generation fighting you for calling it learning and coming up with yet another new term that everyone should use.