Google’s AI thinks I left a Gatorade bottle on the moon
edwardbenson.com
edwardbenson.com
I had a go at something a bit more ambitious a few weeks ago.
If you ask Google Gemini "what was the name of the young whale that hung out in pillar point harbor?" it will tell you that the whale was called "Teresa T".
Here's why: https://simonwillison.net/2024/Sep/8/teresa-t-whale-pillar-p...
(Gemini used to just say "Teresa T", but when I tried just now it spoiled the effect a bit by crediting me as the person who suggested the name.)
1) https://gemini.google.com/ - this one just searches in Google with current language/region/safe-browsing settings and personal adjustments and rewrites top search results as an answer. Generative capabilities are basically not used.
2) https://aistudio.google.com/ - here you can select specific version and generate response with LLM. Retrieval Augmented Generation (i. e. Google Search) is not used.
I suppose you used #1, that's why you cave the correct result. #2 fails. There is a huge group of question where you can immediately find the answer, but LLM struggles. Another question (as example) is "What was the intended purpose of the TORIFUNE satellite in The Touhou Project?".
OpenAI has something similar, providing https://www.bing.com/chat for RAG and https://chat.openai.com for an actual LLM.
One of the drafts had a little more: "Teresa T is the name of the young humpback whale that was spotted in Pillar Point Harbor. She made headlines in September 2024 when she was seen swimming near the shore, drawing crowds and causing excitement among local residents."
```The young whale that visited Pillar Point Harbor in 2024 was named Teresa T.
It was a humpback whale that ventured into the harbor, likely by accident.```
Except they're not people, and they're not actually engaging with anything. It's all literal bullshit.
>It's all literal bullshit.
bahaha. I thought i was cynical...
One of these projects has to change their name!
NotebookLLM was set up two days ago, presumably by "entrepreneurs" eager to monetize all the free fun people have been having with podcast generation in NotebookLM.
The .net version is really poor quality by comparison
And he was ON IT. Like, he ran to his room and grabbed a pencil and paper and put down an essay (okay about 6 or so sentences) about Minecraft, had me type them in, and ran the Notebook, and now he's just showing off both to EVERYONE.
(Yes, he understands it's not real people.)
maybe because they don’t have enough prior experience to compare “the new way” with any given old way.
I imagine that an AI-native generation will be the same, but with respect to AI.
I think you have the age range of people who wrote their own MySpace pages versus those in pre-structured gardens like Facebook slightly wrong.
I figured you’d at least need to study the same group for 10 years as they grow up to really tell.
Obviously a 12 year old might not understand the limitations of a technology, but give that 12 year old 10 years of living with it and they’d be better than their parents.
Right now there are millions of high school students tweaking and testing different inputs against LLMs for very real consequences: their grades.
Meanwhile I barely trust LLMs enough to write a relatively inconsequential piece of code.
If this version of AI is the real deal, the kids that are really depending on it are going to figure out the breakthroughs, not me.
Every technological revolution has tradeoffs. Luckily once the people who knew what we lost finally die off, the complaints will stop and everyone will think the new normal is fine and better.
Buddha may have described the concept of enlightenment but not specifically how to get there.
Which do you think we’re likely to fall into?
Every time we change something “for the better”, we ought to keep in mind that the old way was a solution to some problem that we no longer know or remember.
Oh I have. It was so cool. The R&D on that technology can't happen fast enough.
I'm just lampooning tech apologist tropes.
"I mean, what's not to like about the new normal?"
"Yeah! It's both new *and* better!"Also, I believe serving content different to the Google bot than normal users see absolutely trashes your search ratings.
"You gotta be good. You gotta be top notch."
"It's like he knew what every team needs before he even applied."
Man, oh man. Comedy gold.
If you can dodge a wrench, you can dodge a ball.
Sounds like a (unironic) linkedin post
The algorithmic overlords have long favored "trends" or more seriously content regurgitation. At first it was "have to post something about $topic". Then it was reaction videos. Arguably negative value add content. Then it was all fed back to algorithmic content regurgitators (LLMs) which flood the internet.
The beauty of this recording is that it sounds convincingly like a podcast. It has the podcast-style pacing, over the top praise for the most mundane things. It highlights how narrow is the mean this "content" has regressed to.
It's comedy. Comedy in absurd, but still comedy.
Now that you mention it, that does fit what has at this point become the primary podcast style so I guess it's actually being surprisingly realistic because the thing it tries to mimic is already so artificial.
It's form over substance all over the place. I absolutely love this TED talk: https://www.youtube.com/watch?v=8S0FDjFBj8o
The more pseudo-intellectual consumerism infiltrates our collective psyche, the more the substance becomes irrelevant. Nuance requires thinking. "For every complex problem there is an answer that is clear, simple, and wrong" -- HL Mencken. But knowing this answer makes you feel smart.
I also haven't seen another country (in Europe at least) where politicians across party lines so frequently emphasize in so many ways how great their country is - not even in a jingoistic way, just as a shared cultural consensus.
2. To me personally, it’s extra funny because it’s two people breathlessly discussing.. me.
3. The strange turns of phrase (“That’s Masto”)
4. The stuff it makes up. I’ve never touched the guitar, but apparently I “shred”.
It has some of this quality.
But it's just a super positive spin on the most mundane of topics. There is an emotional play here that you would not normally see in a "resume".
It's like the wrong emotional subcarrier on the topic that is jaring...
It's fascinating. This tech can extract a template for a typical podcast, extrapolate from a mundane CV, plug that to the template and produce a podcast script that your typical copywriter would.
> But it lacks any insight, not even a rehashed conclusion, and doesn’t really seem to integrate the knowledge
Is it the GPT that is lacking here or the source material it learned on converges to this?
The whole point of grand innovations is that they took years of focus on something not very likely.
Like the iPhone. In the 90s could you imagine electronics that literally everyone had in their pocket with _almost no buttons_. Or in the 70s, could u imagine everyone having their own personal computer?
Even in star trek, communicators had buttons.
"Talk about communication skills!"
It’s so good. I’d honestly listen to them talk about your career for 5 episodes lmao
(Not goalpost moving, I certainly think this is impressive.)
Some podcasters actually do this. For example, I've noticed it in some science podcasts where the goal is to make the audience feel like "gee whiz that's an interesting fact." The podcaster will act super surprised to set the emotional tone, but of course they often already knew that fact and will follow up with more detail in a less surprised tone.
That doesn't mean this isn't a bug. But stuff like that reminds me that LLMs may not learn to be like Data from Star Trek. They may learn to be like Billy Mays, amped up and pretending to be excited about whatever they're talking about.
Some podcasts explicitly avoid this by only having a single host do research so the other host can give genuine reactions. E.g. "You're Wrong About" and "If Books Could Kill".
It doesn't mean it's not impressive, but it's hard to describe what isn't realistic about something until you see it. A technology getting things 90% right may still be wrong enough to be noticeable to people, but it's not like you could predict what the 10% that's wrong will be until you try it, and competing technologies may not have the same 10% that's wrong.
You can take the open source models and fine tune them to take on any persona you want. A lot like what the Flux community doing with the Boring Reality fine tune.
We are all enjoying/noticing some repeatable wack behavior of LLMs, but we are seeing the dual wack of humans revealed too.
Massive gains in neural type models and abilities A, B, C, ..., I, J, K, in very little time.
Lots of humans: It's not impressive because can't L, M, yet.
They say people model change as linear, even when it is exponential. But I think a lot of people judge the latest thing as if it somehow became a constant. As if there hasn't been a succession of big leaps, and that they don't strongly imply that more leaps will follow quickly.
Also, when you know before listening that a new artifact was created by a machine, it is easy to identify faults and "conclude" the machine's output was clearly identifiable. But that's pre-informed hindsight. If anyone heard this podcast in the context of The Onion, it would sound perfectly human. Intentionally hilarious, corny, etc. But it wouldn't give itself away as generated.
OpenAI considers this phenomenon a software bug
I think that's what makes notebookllm so realistic. To me this is my perception of all podcasts
With LLMs is far cheaper and easier, but the principle is the very same: trust vs verification or the possibility thereof.
Perhaps in 3-5 years a fully generated influencers by voice and "body" become a thing.
Who wants this shit? I do not want puritanical American corporations telling me what I can and can't use their automated tools for. There's nothing harmful in me performing a computer analysis about Junko Furuta, no more so than Pecco Bagnaia. How have we let them treat us like toddlers? It's infantilising and I won't take part in it. Google, OpenAI, Microsoft, Apple, Meta and the rest of them can shove these crappy "AI" tools.
I’ve had them be incorrect a few times when feeding in arxiv papers, but I don’t think the audience for podcasts like that care.
Uncanny valley all round.
No, NotebookLM creates summaries and podcasts, or answers questions specifically from the documents you feed it.
Feed it fiction it will create fiction as would a human tasked to do the same.
A human might say "are you sure" and understand what they asked and the answer.
If they can do it for fun, malicious people are probably already doing it to manipulate ai answers. Can you imagine poisoning ai dataset with your blackhat SEO work?
If the article had got Gemini AI to tell other users he’d left Gatorade on the moon that would be notable, but this is literally just summarising the document it was given. Usually Google search crawler is fairly good at finding when it has been fed different information and ignores/downgrades the site after a few days/weeks
I'm not sure if it's deliberately deceptive or just an example of poor writing conveying something other than what the author intended, but the attack in the article is not instantiated in the blog post.
Mind you, I well believe that less extreme examples of the attack are possible. However, I doubt truly poisoning an LLM with something that improbable is that easy, on the grounds that plenty of that sort of thing already litters the internet and the process of creating an LLM already has to deal with that. I don't think AI researchers are so dim that they've not considered the possibility that there might be, just might be, some pages on the Internet with truly ludicrous claims on them. That's... not really news.
Ascribing mental qualities to machines poses several challenges. Ethically, it blurs the line between human and machine, raising questions about rights and responsibilities1. Philosophically, it complicates the understanding of mind by attributing human-like qualities to non-human entities23. Practically, it can lead to misunderstandings about the capabilities and limitations of machines, as they do not truly possess beliefs or intentions like humans do56. Additionally, this practice can result in a misuse of language, potentially misleading people about the nature of artificial intelligence
And lastly from http://www.cse.msu.edu/~cse841/papers/McCarthy.pdf
We must be careful to not ascribe mental qualities to machines that do not have that.
Thinking may involve finding minimal/maxima, but it's not a 1-to-1 relation. I'd argue that thinking requires a will component: a sunflower is not a thinking entity because it doesn't have the choice not to follow the sun.
[1] https://physics.stackexchange.com/questions/444307/does-a-ro...
Artificial artificial.
Great little project, though. And, as satire, I did like the show notes writing.
And the generative AI was impressive, in a way. Though I haven't yet thought of a positive application for it. And I don't know the provenance of the training data.
I fail to see the value in doing this.
"Oh hey everybody! I set up a website which presents different content to a crawler than to a human ..... and the crawler indexed it!!"
I'm reading this less as
> "We serve different data to Google when they are crawling and users who actually visit the page"
and more
> "We serve the user different data if they access the page through AI (NotebookLM in this case) vs. when they visit the page in their browser".
The former just affects page rankings, which had primarily interfaced with the user through keywords and search terms -- you could hijack the search terms and related words that Google associated with your page and make it preferable on searched (i.e. SEO).
The latter though is providing different content on access method. That sort of situation isn't new (you could serve different content to Windows vs. Mac, FireFox vs. Chrome, etc.), but it's done in a way that feels a little more sinister -- I get 2 different sets of information, and I'm not even sure if I did because the AI information is obfuscated by the AI processes. I guess I could make a browser plugin to download the page as I see it and upload it to NotebookLM, subverting it's normal retrieval process of reaching out to the internet itself.
Please don’t do this. You don’t need a professional mic to record a podcast with your kids any phone or computer mic will work. Then you can have fun editing it with open source audio tools.
Don’t have a computer generate crap for your kids to consume. Make it with them instead.
Would we have taken the time to compose such a track otherwise? No way. But it's sure been fun playing with what we can. The end result is us laughing together, and I love it.
If my kids are interested, it will spark the path towards "proper" music production where they learn the necessary underlying skills.
Chances are, this won't be the thing they latch onto. They have so much exposure to other things that this doesn't really need to be the hill I die on. They'll find something else to invest in.
Don't (1) ask a machine to automate the creativity and (2) then give it to your kids to consume in a non-interactive manner.
My argument is my kids would never even attempt to do this without the machine's involvement. Making original music the "legacy" way involves days/weeks/months/years of learning. It's just not realistic to argue that they should make music from scratch.
Kids making something, with any tool, is better than consuming something. Because making something is creative. They don’t need to know music theory to have fun making things, it doesn’t even need to be good. The AI tool should allow a kids creativity to go further, not replace it.
The original post was about a parent using an LLM to write a podcast to entertain their kids, which is just another way for kids to passively consume low-quality crap. The proposed alternative was to make a “podcast” with your kid, using any tool you have available at home.
The point was don’t automate the creativity and leave your kids to consume passively.
AI gen is probably the future of music composition. By the time your kids are professionals AI is going to be a lot stronger than it is today.
Are your kids having fun? Are they learning? Is it a good bonding experience? Those are the things that matter.
That said, I think there has been many sci-fi stories written from the perspective of it being true. One day we may face that reality, and we'll have to ask ourselves what parenthood means.
Suffice to say, raising a child is a uniquely wonderful opportunity, for those who choose to embark on that adventure.
This is just the NotebookLM crawler that is being tricked, which is still in it's experimental stage. Rest assured as it scales Google can easily implement safeguards against all spammy tricks people use