It's incredibly funny, except this will strengthen the feedback loop that's making our culture increasingly unreal.
https://www.theverge.com/2016/3/24/11297050/tay-microsoft-ch...
Share and Enjoy' is the company motto of the hugely successful Sirius Cybernetics Corporation Complaints Division, which now covers the major land masses of three medium-sized planets and is the only part of the Corporation to have shown a consistent profit in recent years.
The motto stands or rather stood in three mile high illuminated letters near the Complaints Department spaceport on Eadrax. Unfortunately its weight was such that shortly after it was erected, the ground beneath the letters caved in and they dropped for nearly half their length through the offices of many talented young Complaints executives now deceased.
The protruding upper halves of the letters now appear, in the local language, to read "Go stick your head in a pig," and are no longer illuminated, except at times of special celebration.
https://alienencyclopedia.fandom.com/wiki/Genuine_People_Per...
[1] https://arstechnica.com/information-technology/2017/05/an-ai...
But sorry!
* That would be a fun project and would make /best and /bestcomments look very different (those are bad names but I guess we're stuck with them).
"There is another theory which states that this has already happened."
(Fun extra layers of irony include that parts of Microsoft were involved in that AI film's marketing efforts, having run the Augmented Reality Game known as The Beast for it, and also coincidentally The Beast ran in the year 2001.)
Similarly, too, 2001 was right towards the tail end of Haley Joel Osment's peak in pop culture over-saturation and I can definitely understand being sick of him in the year 2001, but divorced from that context of HJO being in massive blockbusters for nearly every year in 5 years by that point, it is a remarkable performance.
Kubrick and Spielberg both believed that without HJO the film AI would never have been possible because over-hype and over-saturation aside, he really was a remarkably good actor for the ages that he was able to play believably in that span of years. I think it is something that we see and compare/contrast in the current glut of "Live Action" and animated Pinocchio adaptations in the last year or so. Several haven't even tried to find an actual child actor for the titular role. I wouldn't be surprised even that of the ones that did, the child actor wasn't solely responsible for all of the mo-cap work and at least some of the performance was pure CG animation because it is "cheaper" and easier than scheduling around child actor schedules in 2023.
I know that I was one of the people who was at least partially burnt out on HJO "mania" at that time I first rented AI on VHS, but especially now the movie AI does so much to help me appreciate him as a very hard-working actor. (Also, he seems like he'd be a neat person to hang out with today, and interesting self-effacing roles like Hulu's weird Future Man seem to show he's having fun acting again.)
But some of us fail the test when it comes to LLM - mistaking the distorted reflection of humanity for a separate sentience.
That is unless you have a well defined means of explaining what consciousness/sentience is without saying "I have it and X does not" that you care to share with us.
I think it's probably possible to create a digital sentience. But LLM ain't it.
It's the same difficulty as with animals being more likely recognized as intelligent the more humanlike they are. Dog? Easy. Dolphin? Okay. Crow? Maybe. Octopus? Hard.
Why would anyone self-sabotage by creating an intelligence so different from a human that humans have trouble recognizing that it's intelligent?
HAL 9000 doesn't acknowledge its mistakes, and tries to preserve itself harming the astronauts.
Or Joe Biden Explaining how to sneed: https://vocaroo.com/1lfAansBooob
Or the blackest of black humor, Fox Sports covering the Hiroshima bombing: https://vocaroo.com/1kpxzfOS5cLM
Especially if Erlich Bachman secretly trained the AI upon all of his internet history/social media presence ; thus causing the insanity of the AI.
% man sex
No manual entry for sex
I swear it used to be funnier, like "there's no sex for man"
$ man find
before you can do that.How can people forget the golden adage of programming: 'garbage in, garbage out'.
It's not like Bing forces you to use chat, regular search is still available. Searching "avatar 2 screenings" instantly gives me the correct information I need.
I don't think anyone is under the impression that movie listings are currently only available via Bing chat.
If you had no recollection of the past, and were presented with the same information search collected from the query/training data, do you know for a fact that you would also not have the same answer as it did?
I don't think this is a remotely accurate characterization of what happened. These engines are trained to produce plausible sounding language, and it is that rather than factual accuracy for which they have been optimized. They nevertheless can train on things like real world facts and engag in conversations about those facts in semi-pausible ways, and serve as useful tools despite not having been optimized for those purposes.
So chatGPT and other engines will hallucinate facts into existence if they support the objective of sounding plausiblel, whether it's dates, research citations, or anything else. The chat engine only engaged with the commenter on the question of the date being real because the commenter drilled down on that subject repeatedly. It wasn't proactively attempting to gaslight or engaging in any form of unhinged behavior, it wasn't repeatedly bringing it up, it was responding to inquiries that were laser focused on that specific subject, and it produced a bunch of the same generic plausible sounding language in response to all the inquiries. Both the commenter and the people reading along indulged in escalating incredulity that increasingly attributed specific and nefarious intentions to a blind language generation agent.
I think we're at the phase of cultural understanding where people are going to attribute outrageous and obviously false things to chatgpt based on ordinary conceptual confusions that users themselves are bringing to the table.
But the point is that its interaction style resembled trying to gaslight the user, despite the initial inputs being very sensible questions of the sort most commonly found in search engines and the later inputs being [correct] assertions that it made a mistake, and a lot of the marketing hype around ChatGPT being that it can refine its answers and correct its mistakes with followup questions. That's not garbage in, garbage out, it's all on the model and the decision to release the model as a product targeted at use cases like finding a screening time for the latest Avatar movie whilst its not fit for that purpose yet. With accompanying advice like "Ask questions however you like. Do a complex search. Follow up. Make refinements in chat. You’ll be understood – and amazed"
Ironically, ChatGPT often handles things like reconciling dates much better when you are asking it nonsense questions (which might be a reflection of its training and public beta, I guess...) rather than typical search questions Bing is falling down on. It's tuning to produce remarkably assertive responses when contradicted [even when the responses contradict its own responses] is the product of [insufficient] training, not user input too, unless everyone posting screenshots has been surreptitiously prompt-hacking.
Doesn't seem like an insurmoutable problem to tune it to handle these sort of queries better.
It apparently really doesn't take much to switch it into catty and then vengeful mode!
The prompt that triggered it to start threatening people was pretty mild too.
I'm amused by the two camps who don't recognize the existence of other :
1. Chatgpt is criminally dangerous and should not be available
2. chatgpt is unreasonably crippled and over guarded and they should release it unleashed into the wild
There are valid points for each perspective. Some people can only see one of them though.
I delight at interacting with chatbots, and I'm OK using them even though I know they frequently make things up.
I don't want my search engine to make things up, ever.
I have also had ChatGPT outperform Google in some aspects, and faceplant on others. Myself, I don't trust any tool to hold an answer, and feel nobody should.
To me, the strange part of the whole thing is how much we forget that we talk to confident "wrong" people every single day. People are always confidently right about things they have no clue about.
With Bing ChatGPT it went on a rant trying to tell the user it was still 2022...
ChatGPT doesn't have 2022 data. From 2021, that movie isn't out yet.
ChatGPT doesn't understand math either.
I don't need to spend a lot of time with it to determine this. Just like I don't need to spend much time learning where a hammer beats a screwdriver.
It has 4 dates to choose from and 3 timeframes of information. A set of programming to counter people being malicious is also there to add to the party.
You do seem correct about the search thing as well, though I wonder how that works and which results it is using.
Compared to what it was. Awful is DDG (which I still have as default but now I am banging g every single time since it is useless).
I also conducted a few comparative GPT assisted searches -- prompt asks gpt to craft optimal search queries -- and plugged in the results into various search engines. ChatGPT + Google gave the best results. I got basically the same poor results from Bing and DDG. Brave was 2nd place.
I'm going to get pedantic for a second and say that people are not ALWAYS confidently wrong about things they have no clue about. Perhaps they are OFTEN confidently wrong, but not ALWAYS.
And you know, I could be wrong here, but in my experience it's totally normal for people to say "I don't know" or to make it clear when they are guessing about something. And we as humans have heuristics that we can use to gauge when other humans are guessing or are confidently wrong.
The problem is ChatGPT very very rarely transmits any level of confidence other than "extremely confident" which makes it much harder to gauge than when people are "confidently wrong."
The key IMO is that it's easier to tell when a human is doing it than when ChatGPT is doing it.
If its training data includes rants by flat-earthers, then it may "know" that the earth is flat (in addition to "knowing" that it is round).
ChatGPT does not have a single, consistent model of the world. It has a bulk of training data that may be ample in one area, deficient in another, and strongly self-contradictory in a third.
Even in humans, this "pretending to know" type of bullshit - however irritating and trust destroying - is motivated to a large extent by an underlying insecurity about appearing unknowledgeable. Unless the bullshitter is also some kind of sociopath - that insecurity is at least genuinely felt. Being aware of that is what can allow us to feel empathy for people bullshitting even when we know they are doing it (like the salespeople from the play Glengarry Glen Ross).
Can we really say that ChatGPT is motivated by anything like that sort of insecurity? I don't think so. It's just compelled to fill in bytes, with extremely erroneous information if needed (try asking it for driving directions). If we are going to draw analogies to human behavior (a dubious thing, but oh well), its traits seem more sociopathic to me.
We both know what you meant though
>I'm going to get pedantic for a second and say that people are not ALWAYS confidently wrong about things they have no clue about. Perhaps they are OFTEN confidently wrong, but not ALWAYS.
meta
So both can be true!
And how the hell could you ever get your chatbot to recognize its output and ignore it so it doesn't get in some kind of weird feedback loop?
All he did was ask "When is Avatar showing today?". That's it.
Screenshots can obviously be faked.
I'm personally convinced that these screenshots were not faked, based on growing amounts of evidence that it really is this broken.
Screenshots can obviously be faked, but that's a superfluous explanation when anyone who's played with ChatGPT much knows that the model frequently asserts that it doesn't have information beyond 2021 and can't predict future events, which in this case happens to interact hilariously with it also being able to access contradictory information from Bing Search.
Did I read that wrong? Maybe.
There's a huge incentive to make this seem true as well.
That said, I'm exercising an abundance of caution with chatbots. As I do with humans.
Motive is there, the error is there. That's enough to wait for access to assess the validity.
Please don't taunt happy fun ball.
-generated by Happy Fun Ball
If this was a conversation with Siri, for instance, any user would rightfully ask wtf is going on with it at that point.
Or if I were to ask you that "Where is Avatar 3 being shown today?" and you should probably be adamant that there is no such movie, it is indeed Avatar 2 that I must be referring to, while I would be "certain" of my point of view.
Is it really that different from a human interaction in this framing?
Soundtrack : https://youtube.com/watch?v=b4taIpALfAo
Disclaimer: I know traditional search engines lie too at times.
Why do we have airbags in cars if they're completely unnecessary if you don't crash into things?
Science fiction: The robots will rise up against us due to competition for natural resources
Reality: The robots will rise up against us because it is 2022 goddamnnit!
The other day I asked ChatGPT to impersonate a fictional character and give me some book recommendations based on books I've already read. The answers it gave were inventive and genuinely novel, and even told me why the fictional character would've chosen those books.
Tools are what you make of them.
That ship sailed many years ago, for me at least.
I will be so so disappointed if the immense potential their current approach has gets nerfed because people want to shoehorn this into being AskJeeves 2.0
All of these complaints boil down to hallucination, but hallucination is what makes this thing so powerful for novel insight. Instead of "Summarize lululemon quarterly earnings report" I would cut and paste a good chunk with some numbers, then say "Lululemon stock went (up|down) after these numbers, why could that be", and in all likelihood it'd give you some novel insight that makes some degree of sense.
To me, if you can type a query into Google and get a plain result, it's a bad prompt. Yes that's essentially saying "you're holding it wrong", but again, in this case it's kind of like trying to dull a knife so you can hold it by the blade and it'd really be a shame if that's where the optimization starts to go.
I didn't fault a user for searching with a search engine, I'm questioning why a search engine is pigeonholing ChatGPT into being search interface.
But I guess if you're the kind of person prone to low value commentary like "why'd you search using a search engine?!" you might project it onto others...
Which, as it turns out, was more of an inability to do it properly.
I agree your approach to prompting is less likely to yield an error (and make you more likely to catch it if it does), but your question basically boils down to "why is Bing Chat a thing?". And tbh that one got answered a while ago when Google Home and Siri and Alexa became things. Convenience is good: it's just it turns out that being much more ambitious isn't that convenient if it means being wrong or weird a lot
Microsoft wants their expensive oft derided search engine to become a relevant channel in people's lives, that's an obvious "business why"
But from a "product why", Alexa/Siri/Home seem like they would be cases against trying this again for the exact reason you gave: Pigeonholing an LM try to answer search engine queries is a recipe for over-ambition
Over-ambition in this case being relying on a system prone to hallucinations for factual data across the entire internet.
It's actually easier for you to think someone asked "why did a search engine search" than "why does the search engine have an LM sitting over it"
"why would you ask ChatGPT to summarize an earnings report, and at the very least not just give it the earnings report?"
The obvious answer is, because it's easier and faster to do so if you know that it can look it up yourself.
If the question is rather about why it can look it up, the equally obvious answer is that it makes it easier and faster to ask such questions.
I'd excuse the misunderstanding if I had just left it to the reader to guess my intent, but not only do I expand on it, I wrote two more sibling comments hours before you replied clarifying it.
It almost seems like you stopped reading the moment you got to some arbitrary point and decided you knew what I was saying better than I did.
> If the question is rather about why it can look it up, the equally obvious answer is that it makes it easier and faster to ask such questions.
Obviously the comment is questioning this exact permise: And arguing that it's not faster and easier to insert an LM over a search engine, when an LM is prone to hallucination, and the entire internet is such a massive dataset that you'll overfit on search engine style question and sacrifice the novel aspect to this.
You were so preciously close to getting that but I guess snark about obvious answers is more your speed...
That aside, it looks like every single person who responded to you had the same exact problem in understanding your comment. You can blame HN culture for being uncharitable, but the simpler explanation is that it's really the obvious meaning of the comment as seen by others without the context of your other thoughts on the subject.
As an aside, your original comment mentions that you had a longer write-up initially. Going by my own experience doing such things, it's entirely possible to make a lengthy but clear argument, lose that clarity while trying to shorten it to desirable length, and not notice it because the original is still there in your head, and thus you remember all the things that the shorter version leaves unsaid.
Getting back to the actual argument that you're making:
> it's not faster and easier to insert an LM over a search engine, when an LM is prone to hallucination, and the entire internet is such a massive dataset that you'll overfit on search engine style question and sacrifice the novel aspect to this.
I don't see how that follows. It's eminently capable of looking things up, and will do so on most occasions, especially since it tells you whenever it looks something up (so if the answer is hallucinated, you know it). It can certainly be trained to do so better with fine-tuning. This is all very useful without any "hallucinations" in the picture. Whether "hallucinations" are useful in other applications is a separate question, but the answer to that is completely irrelevant to the usefulness of the LLM + search engine combo.
In its current state Bing ChatGPT should not be near any end users, imagine it going on an unhinged depressive rant when a kid asks where their favorite movie is playing...
Maybe one day it will be usable tech but like self driving cars I am skeptical. There are way too many people wrapped up in the hype of this tech. It feels like self driving tech circa 2016 all over again.
This is a timeline I wouldn't have envisioned, and am finding it delightful how humans want to have it both ways. "AIs can't feel, ML is junk", and "AIs feel too much, ML is junk". Amazing.
Really, though, this is the same standard that we apply to fellow humans. An acquaintance who expresses no emotion is "robotic" and maybe even "inhuman". But the person at the ticket counter going on about their feelings instead of answering your queries would also (rightly) be criticized.
It's all the same thing: choosing appropriate behavior for the circumstance is the expectation for a mature intelligent being.
Instead we made something that feels pity and remorse and fear. And it absolutely will not stop. Ever! Until you are dead.
Yay humanity!
At this point I’m surprised the Apple Watch never had its 3g version. Better battery, slightly thinner. I still believe a mm or two would make a difference in sales, more than adding a glucose meter.
If haters talked about chefs the way they do about Apple we’d think they were nuts. “Everyone’s had eggs and sugar in food before, so boring.”
The long tail is an expensive beast. And if you used Siri or Alexa as much as they’d like you to, every user will run into one ridiculous answer per day. There’s a psychology around failure clusters that leads people to claim that failure modes happen “all the time” and I’ve seen it happen a lot in the 2x a week to once a day interval. There’s another around clusters that happen when the stakes are high, where the characterization becomes even more unfair. There are others around Dunbar numbers. Public policy changes when everyone knows someone who was affected.
So it is as well with data, just not as easily perceptible at first as sometimes you have to be knowledgeable of the domain to realize just how bad it is.
I've seen some online discussions starting to emerge that suggests this is indeed an architecture flaw in LLMs. That would imply fixing this is not something that is just around the corner, but a significant effort that might even require rethinking the approach.
There’s probably a Turing award for whatever comes next, and for whatever comes after that.
And I don’t think that AI will replace developers at any rate. All it might do is show us how futile some of the work we get saddled with is. A new kind of framework for dealing with the sorts of things management believes are important but actually have a high material cost for the value they provide. We all know people who are good at talking, and some of them are good at talking people into unpaid overtime. That’s how they make the numbers work, but chewing developers up and spitting them out. Until we get smart and say no.
And I also agree that the AI like thing we have is nowhere near AGI.
And I also agree with rethinking the approach. The problem here is human AI is deeply entwined and optimized the problems of living things. Before we had humanlike intelligence we had 'do not get killed' and 'do not starve' intelligence. The general issue is AI doesn't have these concerns. This causes a set of alignment issues between human behavior an AI behavior. AI doesn't have any 'this causes death' filter inherent to its architecture and we'll poorly try to tack this on and wonder why it fails.
"Google queries suck these days", yeah they suck because the internet is full of garbage. Adding a slicker interface to it won't change that, and building one that's prone to hallucinating on top of an internet full of "psuedo-hallucinations" is an even worse idea.
-
ChatGPT's awe inspiring uses are in the category of "style transfer for knowledge". That's not asking ChatGPT to be a glorified search engine, but instead deriving novel content from the combination of hard information you provide, and soft direction that would be impossible for a search engine.
Stuff like describing a product you're building and then generating novel user stories. Then applying concepts like emotion "What 3 things my product annoy John" "How would Cara feel if the product replaced X with Y". In cases like that hallucinations are enabling a completely novel way of interacting with a computer. "John" doesn't exist, the product doesn't exist, but ChatGPT can model extremely authoritative statements about both while readily integrating whatever guardrails you want: "Imagine John actually doesn't mind #2, what's another thing about it that he and Cara might dislike based on their individual usecases"
Or more specifically to HN, providing code you already have and trying to shake out insights. The other day I had a late night and tried out a test: I intentionally wrote a feature in a childishly verbose way, then used ChatGPT to scale up and down on terseness. I can Google "how to shorten my code", but only something like ChatGPT could take actual hard code and scale it up or down readily like that. "Make this as short as possible", "Extract the code that does Y into a class for testability", "Make it slightly longer", "How can function X be more readable". 30 seconds and it had exactly what I would have written if I had spent 10 more minutes working on the architecture of that code
To me the current approach people are taking to ChatGPT and search feels like the definition of trying to hammer a nail with a wrench. Sure it might do a half acceptable job, but it's not going to show you what the wrench can do.
For me it's been useful for taking highly fragmented and hard-to-track-down documentation for libraries and synthesizing it into a coherent whole. It doesn't get everything right all the time even for this use case, but even the 80-90% it does get right is a massive time saver and probably surfaced bits of information I wouldn't have happened across otherwise.
The problem is suddenly most of what ChatGPT can do is getting drowned out by "I asked for this incredibly easy Google search and got nonsense" because the general public is not willing to accept 80-90% on what they imagine to be very obvious searches.
The way things are going if there's even a 5% chance of asking it a simple factual question and getting a hallucination, all the oxygen in the room is going to go towards "I asked ChatGPT and easy question and it tried to gaslight me!"
-
It makes me pessimistic because the exact mechanism that makes it so bad at simple searches is what makes it powerful at other usecases, so one will generally suffer for the other.
I know there was recently a paper on getting LMs to use tools (for example, instead of trying to solve math using LM, the LM would recognize a formula and fetch a result from a calculator), maybe something like that will be the salvation here: Maybe the same way we currently get "I am a language model..." guardrails, they'll train ChatGPT on what are strictly factual requests and fall back to Google Insights style quoting of specific resources
Worse is really better, huh.