AI Hype is completely out of control – especially since ChatGPT-4o [video]
youtube.com
youtube.com
The problem is that capital isn't that patient. People are sinking billions into LLM integrations, at discount rates of 5%+. If it takes 20 years for a tech to pay off, at a 5% discount rate, and you sunk a billion dollars into it, it needs to earn back $2.65B. Moreover, at some point during that 20-year period, somebody's going to ask "Where's my billion dollars?" and pull the plug on the project.
I think the tech behind LLMs may eventually be game-changing, but that tech is going to change ownership and get reinvented several times, and we won't actually see profitable, sustainable benefits until industries get refactored into cost structures that make sense for what LLMs can actually do. The web needed its Netscape and Yahoo and Geocities and Apache, but it also needed Google and Rails and Django and GitHub and Stripe and Facebook and Webkit and nginx and MySQL to really become what it did.
You should see them as a passage - something in study and development today for something more complete tomorrow.
It also does concern me how basically nobody is building a real product around the current state of "AI", but are rather hinging their success on what they believe it will be in the near future.
This morning I met an newspaper article: "What will be the results of the European elections? Let us ask ChatGPT".
But see, being it a given from experience that humans can be so bad in judgement, reasoning, professionalism, output... Why did we strive for superhuman judgement, reasoning, professionalism, output? (The same way we strived for superhuman strength.)
> Why did we strive for superhuman judgement, reasoning, professionalism, output?
Lots of possible reasons; there's many ways for it to be valuable.
(And it's not like the newspapers are deliberately writing fiction, with notable exceptions like The Onion).
If it's a rhetorical question, I'd be interested to know what you had in mind :)
edit: to be clear, I'm saying that we are running into scaling "walls" that make hard extrapolation based on increased investment senseless.
[1]https://www.gartner.com/en/articles/what-s-new-in-the-2023-g...
It will likely be disruptive in some areas short term, but will take much longer in other areas.
What complex applications are you speaking of?
The past year has brought us both model improvements along with drastic cost reductions. Its been a pretty magical year imo.
Is it AGI? Not even close but we have been utilizing the tooling improvements to build products internally.
It's one of those short term is overhyped, long term is underestimated things.
Not too many applications can justify spending huge amounts of energy generating answers that may just be nice sounding BS.
We may finally see people realising that just because they can use a hammer on a screw, doesn't make it the best choice.
it makes sense because big money is betting it's worth the investment -- but it may not be
There are also plenty of uses for LLMs beyond generating hopefully-accurate answers, such as for fictional content or use as foundation models for tasks like translation. Though we are definitely in the "throw things at the wall and see what sticks" stage currently.
* User interfaces. It does provide a genuinely new way to interface seamlessly with existing software.
* Translation.
* Generating bullshit. Yes, there's demand for this even if perhaps there shouldnt be.
* Not much else. That includes using it as a specialized autocomplete. I think it falls down pretty badly at that.
Not so good at anything that's (for want of a better term) "creative".
It's great as a non offensive stack overflow replacement, but just look at aider benchmarks (amazing work by the way): most capable models really struggle to make basic changes to a real code base.
Does anybody actually using it in practice believes the hype? I thought the hype is just another theatre for investors
EDIT: just some points taken from: https://aider.chat/docs/leaderboards/#code-editing-leaderboa...
- The metric "Percent completed correctly" maxes out at 72.9% with gpt4o, while at the same time, giving out correctly formatted output only 96.2% of the time.
- Benchmark suite is based on https://github.com/exercism/python, which very likely is a part of the training material already! In real code bases, no LLM would have the advantage of seeing new or proprietary code
For coding I have long given up on it. I only see it useful for people that want to do small scripts in languages they are not familiar with or to write toy examples of web/mobile apps.
For art I can sort of see a workflow in adding details to textures and using existing images and 3d assets to guide the generation of images. This is all just my speculation so I could be completely wrong just like the person that knowns very little to no programming thinking they will be able to code whatever they want with it.
The only skill I brought to the table was app architecture.
It's not just hype.
I also built multiple things with it and I always came to a point, where it just couldn't handle anything slightly larger than a mvp, or a non guided change requiring editing multiple files at once
The best way to use it is to make everything as atomic as possible. Ask it for one function at a time, rather than "make my app handle user auth"
In the future, you could prompt GPT10 with "give me a marketing plan" and its output would be just as terrible as GPT3's.
Leveling up one's prompting skill from zero shot to few shot to agenetic is how you get usable results.
What do you do, in order for chatgpt to be able to pick up your library and it's patterns? What obstacles I see in this scenario:
* Base models takes months and millions to train
* RLHF supposedly can add knowledge, but it's disputed to mostly "change style"
* What incentive will OpenAI have to include your particular library's documentation?
I imagine, if that library starts being really popular, a lot of other code will include examples how to you use it. What about before that?
Including new knowledge always lags (are there two gpt updates per year? maybe a up to 4, but not really significantly more) few months, so what about a fast moving agile greenfield project? It could cause frustration in LLM users (I know I have been bitten a lot by some python library changes already).
It seems that it's just another tool in the box for humans to use. In far far future maybe, when we somehow get around those millions of dollars for fine tuning (doubtful) and/or libraries simply stop changing.
But still, put any really not small code base into 120k token context and see how easy both gpt and cluade opus trip up on themselves. It's amazing, when it works, but currently it's a roll of a dice still
Disagree with this statement - it's too broad. There are plenty of tasks where AI's can fully outperform humans, and the number of tasks where AI's can fully outperform humans seems to increase every week.
Then it depends what we mean by "human level" - the median human is different to the smartest person in the room.
So IMO if we are saying that "there's no evidence that computers will have human level intelligence" we need to take a step back and define what that means - because if you were sitting in the 1960's and you defined it, a modern LLM may very well have already met that definition.
But I think we have at least moved from a state of computers that are less 'intelligent' than a fish, to more 'intelligent' than a dog (assuming by intelligent we mean 'ability to solve problems, and to apply knowledge to novel situations' - i.e. while GPT-4 will make illegal chess moves, it would make less illegal moves than a well-trained dog).
We have moved that quickly from fish to dog, so we need to be careful to not let hubris make us think we are that much more special! Dog-intelligence to human-intelligence is a big leap, but maybe not as much of a leap as fish to dog? Or at least dog to ape and ape to human?
You mean 1967, right? https://en.wikipedia.org/wiki/ELIZA
> However, many early users were convinced of ELIZA's intelligence and understanding, despite Weizenbaum's insistence to the contrary.
* Real humans 66%
* GPT-4: 49.7%
* ELIZA: 22%
* GPT-3.5: 20%
https://arxiv.org/pdf/2310.20216
(I'm rather surprised by ELIZA beating 3.5, as were the researchers).
Turing's introduction of the test, was a 70% chance of spotting the AI after 5 minutes.
Not all world is "big data".
Mind you, this is also all early gen stuff.
This just seems like putting a human-constraint on AI systems so that we can still classify ourselves as special.
“The question of whether a computer can think is no more interesting than the question of whether a submarine can swim.” ― Edsger W. Dijkstra
I was using it with my daughter to help her with math review for high school finals, and was presented by ChatGPT 4o with an impossible triangle as an example problem for her to solve.
We "told" ChatGPT the problem couldn't be solved with this kind of triangle and asked it to come up with other examples NOT using impossible triangles and it couldn't NOT come up with workable examples.
Moreover, ML has historically been very brittle compared to the generalization humans provide. A machine that reads addresses might be 10x better than a person, but it may not be worth it if you can’t do things like say “we don’t even sort the blue envelopes—there’s a specific rule for those” like you could with a human. (The specific case here is irrelevant—humans can handle arbitrary variation on the rules infinitely better than a bare CNN.)
It’s this ability to cope with human-level ambiguity that has so much potential in LLMs. But even they don’t work like people—giving it a new rule could make it worse at all other tasks with no explanation as to why.
Even ignoring the people who insist it comes with baggage of consciousness or qualia.
AI can train on millions to billions of examples in a matter of months as transistors outpace synapses by the degree to which marathon runners outpace continental drift… but AI also need to do that well just to get up to the current level.
If AI could learn from as few examples as we require, Tesla would have achieved level 5 driving years ago — but as is, they've got millions of vehicles with years of real world experience and it's still merely "ok", not even "expert".
If we never learn how to make machines learn faster, then we'll have an economy of people whose jobs are essentially "give the machines examples to learn from"… and for some of those jobs, they could spend years providing those examples and still no machine will do as well.
Also, expect a backlash to AI taking human service jobs. Like having to scan a QR code at a restaurant. Its a bad experience that shows you're enraptured by a technology and don't care about the customer.
I don't disagree that AI will be disruptive. I just think it's going to find a place in technical fields and more people will want to remain talking and interacting with people like we've done for the last 100K years.
Has it really? I haven't noticed any drastic decline over the years. Maybe some, but I wouldn't describe it as "killed".
What universal ATMs and cards have really killed off is personal checks, IMO.
But ask yourself this, when did you last go into a bank. I've done a home loan application on a smart phone.
Anyway, stuff is changing, There's always going to be jobs about. People are flexible and can learn things, AI ain't all that.
And "absolutely chang[ing] the world" is often assumed to be good, but that's not necessarily so. There are a lot of mediocre humans in the world, only capable of rote work. If we push automation to the maximum, eventually we'll reach a point were those people have nothing economically viable to do.
And then you reach a decision point: support them on welfare, forever? Immiserate them and hope they die off (what our current social framework will invariably choose)?
I disagree with that premise. For certain businesses it's fundamentally true, and right now it's broadly true because of the limitations of our technology, but I don't think it's fundamentally true for capitalism.
If the AGI hype pans out, I think we'd see an transition into an economy with a lot fewer consumers. The capitalists would eventually amass all the valuable resources for themselves, and the positive externality that currently drives the consumer economy (the need for labor to make use of capital) would run down and eventually die. Then you'd have a small cadre of elites who own the automation and the resources, and use them chiefly for vanity projects and maybe some B2B activity between themselves (e.g. electricity sales), a small somewhat larger cadre of "middle class" providing luxury/vanity goods to the core elite (e.g. high end prostitutes, Mars colonists), and a great mass of economically useless people living at the margins.
> Ergo, to keep the wheels turning, either everyone is going to have to become an artisan, or UBI is going to have to be a thing.
So the former won't work because artists can't feed themselves by trading art among themselves, and the latter's not necessary because an Elon Musk focused on personal vanity projects doesn't need consumers, he just needs resources an an AGI bureaucracy with AGI robots to command around.
It is constrained by physical infrastructure like compute and memory, just like the brain which had to become physically bigger to get better.
Human-level intelligence should adapt to many uses.
It may require teachers, but not prompt engineers.
The article author is very valid in the stance that this is far from intelligence.
Yes yes, you wouldn’t criticize a fish’s intelligence for not being able to climb a chessboard.
But these things are stupid as hell. They are not even close.
Firstly, there is absolutely no way anyone could know this. Secondly, no they aren't.
Just as one can say, in response to a perfectly correct answer, "Explain what's wrong with your answer" and it will find something, one can either create a vacuum it feels compelled to fill with hallucinations or ask in a way it stays grounded.
That’s the description of a house of cards rather than some revolutionary new technology.
I used the opportunity to send some queries which GPT-3.5 was bad at, and while GPT4 was better and closer to the correct answers, it still failed to produce them. When I pointed out the problems, it got even closer, but still ended up insisting that its (wrong) answer was correct in the end.
So, yeah, GPT4 is much better than GPT3.5, but it's got a long way to go before it becomes real AI.
Currently, I still use it for 3 cases:
1. When exploring something completely new. It saves time reading and learning because you can ask questions with "layman words" and get to learn the specific terms used in the domain quickly.
2. When google search is swamped by SEO spam or google insists on ignoring some of the keywords
3. When I need some boilerplate code that has been written a thousand times. It saves time which would otherwise get spent on concatenating different code segments from 5 different Stackoverflow threads and reading the docs afterwards to ensure it's correct. Now I just copy/paste from ChatGPT and read the docs afterwards to fix any places where it messed up.
It's not as great a gpt4O, but with gemini and 1M tokens (2M now also avaiable) you can dump in a small library of "writing javascript for X" books before asking it to produce code for X. It dramatically improves output.
ChatGPT also can be front loaded with documents, but the context is much smaller (32K in the web app, I believe?)
I don't disagree. Just wanted to highlight the poor intuition of the guy doing the video.
So rather than watching the video for 22 minutes, I felt I was able to get the gist of it in 1 minute...and denied the creator a full view. Was it accurate? I skimmed the video too in order to see, and it seemed correct, but I am aware of the inaccuracy these models tend to have.
I don't think the hype is necessarily accurate in it's specific predictions, but I do think it is accurate in the "things are going to change dramatically" way.
The RN was describing a new AI service that her practice is using. It is a plugin to their tele-medicine software that listens to an ongoing call and uses speech-to-text plus whatever AI magic sauce the product has to fill out a "call sheet". I don't know precisely what a call sheet is, but the RN said it was the worst part of her job and she was very happy that it would automate the paper-work that takes up a large part of her day.
The marketing manager laughed and said she used GPT multiple times a day. As in, she doesn't even send emails now without running the text through GPT. She then relayed an anecdote about how she had taken a picture of her fridge and asked GPT to identify the food available in there and to make a weekly meal plan for her family.
These anecdotes were just brought up spontaneously. That is, the RN was just excited and musing about a new huge time-saver that was making her work-life better. And the marketing manager, who has already completely integrated GPT into her work, was showing how she was now integrating GPT into her family life.
Even though I understand the "facts based reasoning" that is advocated in this video - my personal experience is seeing excitement not at hype, but excitement by non-tech people in the current uses of the technology. The trend I see in the people I am talking to is that AI, even the weak LLM version we are seeing now, is being integrated into work-life and personal-life at an astonishing pace.
https://www.theonion.com/guy-who-sucks-at-being-a-person-see...
I'm on the verge of canceling my subscription altogether, as trying to get the damned thing to do what I ask is starting to take more time than just doing it by hand.
If anyone from OpenAI is reading this: I (and I suspect many others) would rather pay more money than use a shitty model.
we have a lot of people with a lot of investments that benefit greatly from hyping the products, an investment gold rush
we simultaneously have an industry of people leveraging better jobs, job functions, job titles, by showing off their skills with this new technology making real things very quickly, and causing companies to have fomo and react without planning by overhiring and overpromoting
and then finally we have a third class of people just talking about this new technology for the benefits of proximity. if there's a rumor george clooney is is crashing a party, half of people will make sure they're the loudest voice for the perceived benefits of that proximity. and even if it's not accurate, it's effective
all of it's happening at the same time
So, will they do better? I personally think that knowledge is about implications. Notably, things that are similar do not necessarily have identical implications. And I can't see LLM tackling this issue anytime soon. (I even doubt that LLMs provide a suitable approach to this, at all.) So we'll probably end up with "we managed to get rid of reliability", which may be just the opposite of the role, the computer plays in Star Trek scripts.
*) To provide a bit of context to reliability and repeatability: These are really externalization of the major generalizations of Husserl's life-world (Lebenswelt), as in "and so on" and "I can do this again". We are generally fine with machines, because they embody these principles in confined parameters. However, there's also a certain danger in this, as we tend to generalize and idealize along these lines, as this is essential to how we construct and navigate the world, we live in. We'll also apply this to things that are just "mostly", and are apt to be fooled by this, as may be the case with LLMs. (And in the past, machines were a great way out of this trap, as long – and as soon – as we were able to verify the parameters and boundaries.)
And "intelligence" can be defined. People just think of different things when you use the word unless you define it. Here is how I like to define intelligence: the intersection between knowledge and reason. An example of a system that is highly knowledgeable but lacks intelligence due to its inability to reason is Wikipedia. A system that is highly reasonable but lacks intelligence due to its lack of knowledge is a calculator.
It's obvious that language models encode more knowledge than humans do. Even a few GB model that I can download and run on my computer can spit out facts about random things. Way more than any human could. Sometimes they are wrong about facts, but so is any knowledge system, including humans.
An LLMs reasoning capability is a bit less clear cut. When you have any understanding of how LLMs work, it seems like they should have no ability to do novel reasoning. Yet when you try to push them, they are often surprisingly good. I think that people who say the best LLMs have worse than human reasoning capabilities probably overestimate the average adults reasoning capability. Plus it is not hard at all to hook up a language model as a component in a greater intelligence system, by adding other components that specialize in reason. For example an environment to run code in.
Again, it's not the best tool available for training and logic applications, but the fact that it really can demonstrate novel reasoning along with its obvious knowledge means thar it clearly has some non-zero level of intelligence.
I think intelligence might not just be reasoning but also the ability to seek knowledge and reason. Some form of will to attain higher intelligence. Machines only have human interests of further knowledge in that regard.
I would just describe this as intelligent but less than average human intelligence.
> I think intelligence might not just be reasoning but also the ability to seek knowledge and reason. Some form of will to attain higher intelligence.
This is a circular definition of intelligence that doesn't really make sense. How can intelligence be the drive for more intelligence? Those have to be two different things.
Consider a college educated person working in some field that requires lots of critical thinking. They have been doing their job for 30+ years, and they are very good at it. They are stubborn and set in their ways; not interested in changing anything. Now consider a 3 year old who is curious about everything. Who is more intelligent? The answer seems to obviously be the adult, even though they refuse to learn anything more.
I really don't think the willingness or ability to learn is a component of intelligence. A system can be statically intelligent or it can be dynamically intelligent. Dynamically intelligent systems are more exciting and novel, but statically intelligent systems can still be extremely useful.
Ability to learn is more interesting in the context of general intelligence (also note general intelligence implies the existence of non-general or domain specific intelligence). In order to be able to act intelligently in any situation (such as: what if my computer suddenly grows legs) it must be able to learn through experimentation.
I am not stating that intelligence is the drive for more intelligence. I am saying that intelligence is that act of collecting more knowledge. Setting bars like:
> I would just describe this as intelligent but less than average human intelligence.
I know some people who appear dumber than dogs because they believe after schooling is complete they have everything they need to operate in life. Are these people below human intelligence?
Intelligence is restrictive to your exposure to knowledge. All knowledge can be aggregated to a domain. You could state “general intelligence” is a primary school education, but not all primary schools deliver the same education. So what is it?
1. He reminds viewers of deliberate attempts to take advantage of known human gullibility toward so-called "AI" ("Eliza Effect"), using dark patterns.^FN1
2. He cites various instances where so-called "tech" companies have lied about "AI" products and/or faked "AI" demos.
3. The creator of the video said he has included some sixty links to sources. He asserts that comments about "AI" that do not cite sources are unpersuasive. (That will not stop HN commenters from sharing their unsupported opinions.)
FN1. He tells viewers that "cute" is a dark pattern, e.g., a giggling "AI" voice. Long have I wondered about all the silly non-descriptive or misleading names chosen for contemporary software, for so-called "tech" companies, as well as the silly graphics and mascots. What is their purpose. For example, is the reason for the silly company names chosen by so-called "tech" companies more than just avoiding trademark disputes. (Yes.) If software developers and so-called "tech" companies intentionally use these tactics to trick people into thinking or doing something that is against those peoples' interests, then arguably these could be dark patterns. Obvious example of a misleading name: "OpenAI" is not open, and that may have serious consequences, but it's unlikely the company will be changing its name any time soon. The video creator mentions that the SEC has stated it is cracking down on the use of "AI washing" where companies add "AI" to names or otherwise use terminology to trick people into believing "AI" is used when in fact it is not.
In sum, the video is about the Silicon Valley "culture of lying". Perhaps ironic that the "AI" being pitched by SillyCon Valley today has no concept of a "lie". Despite the endless anthropomorphism, a computer running "AI" has no concepts whatsoever. Concepts come from people, not computers.
Who is left ?
narrative shift triggered. needs to be combated with more hype. VC money is at stake. co-pilots are threatened.
sora-hype is exhausted. 4o-hype is tired.
call-to-action: "thought-leaders" like a16z to put out giant "architecture of ai" articles to keep the fomo alive.