NotebookLM's automatically generated podcasts are surprisingly effective
simonwillison.net
simonwillison.net
I’d feed it the Singularity paper, but I’m not sure I need that extra boost of anxiety these days…
In other words: soapbox is presumably some sort of toy car that goes 15mph, and formula 1 goes up above 150mph at least (as you can tell, I’m not a car guy). If you have any actual scientific argument as to why a model that can score 90-100 on a typical IQ test has only 1/10th the symbolic reasoning skills of a human, I’d love to eat my words! Maybe on some special highly iterative, deliberation-based task?
It's also not at all clear to me what "symbolic" could mean in this context. If it means the software has concepts, my response would be that they aren't concepts of a kind that we can clearly recognize or label as such (edit: and that's to say nothing of the fact that the ability to hold concepts/symbols and understand them as concepts/symbols presupposes internal life and awareness).
The best analogy I've heard for these models is this: you take a completely, perfectly naive, ignorant man, who knows nothing, and you place him in a room, sealed off from everything. He has no knowledge of the outside world, of you, or of what your might want from him. But you slip under the door of his cell pieces of paper containing mathematical or linguistic expressions, and he learns or is somehow induced to do something with them and pass them back. When what he does with them pleases you, you reward him. Each time you do this, you reinforce a behavior.
You repeat this process, over and over. As a result, he develops habits. Training continues, and those habits become more and more precisely fitted to your expectations and intentions.
After enough time and enough training, his habits are so well formed that he seems to know what a sonnet is, how to perform derivatives and integrals, and seems to understand (and be able to explain!) concepts like positive and negative, and friend and foe. He can even write you a rap-battle libretto about nineteenth-century English historiography in the style of Thomas Paine imitating Ikkyu.
Fundamentally, though, he doesn't know what any of these tokens mean. He still doesn't know that there's an outside world. He may have ideas that are guiding his behavior, but you have no way of knowing that -- or of knowing whether they bear any resemblance to concepts or ideas you would recognize.
These models deal with tokens similarly. They don't know what a token is or represents -- or we have no reason to think they do. They're just networks of weights, relationships, and tendencies that, from a given seed and given input, generate an output, just like any program, just like your phone keyboard generates predictions about the next word you'll want to type.
Given billions and billions and billions and billions of parameters, why shouldn't such a program score highly on an IQ test or on the LSAT? Once the number of parameters available to the program reaches a certain threshold (edit: and we've programmed a way for it to connect the dots), shouldn't we be able to design it in such a way that it can compute correct answers to questions that seem to require complex, abstract reasoning, regardless of whether it has the capacity to reason? Or shouldn't we be able to give it enough data that it's able to find the relationships that enable it to simulate/generate patterns indistinguishable from real, actual reasoning?
I don't think one needs to be cynical to be unimpressed. I'm unimpressed simply because these models aren't clearly doing anything new in kind. What they're doing seems to be new, and novel, only because of the scale at which they do what they do.
Edit: Moreover, I'm hostile to the economic forces that produced these models, as everybody should be. They're the purest example of what Jaron Lanier has been warning us about -- namely that, when information is free, the wealthiest are going to be the ones who profit from it and dominate, because they'll be the ones able to pay for the technology that can exploit it.
I have no doubt Altman is aware of this. And I have no doubt that he's little better than Elizabeth Holmes, making ethical compromises and cutting legal corners, secure in the smug knowledge that he'll surely pay his moral debts (and avoid looking at the painting in the attic) and obviously make the world a better place once he has total market dominance.
And none of the other major players are any better.
This is very "humans are just hunks of matter! They can't think!".
More specifically, my comment aims to meet the challenge posed by the person I answered:
> I highly recommend letting that fact be a little bit impressive, someday. There’s no way you live through any event that’s more historically significant, other than perhaps an apocalypse or two. [...] If you have any actual scientific argument as to why a model that can score 90-100 on a typical IQ test has only 1/10th the symbolic reasoning skills of a human, I’d love to eat my words.
I have no idea what would constitute a "scientific argument" in this instance, given that the challenge itself is unscientific, but, regardless, the results that so impress this person are, without question, achievable without reasoning, symbolic or otherwise. To say that the model "muses" or "has [...] symbolic reasoning" is to make a wild, arbitrary leap of faith that the data, and workings of these models, do not support.
The models are token-prediction machines. That's it. They differ not in kind but in scale from the software that generates predictions in our cell-phone keyboards. The person I answered can be as impressed as he wants to be by the high quality he thinks he sees in the predictions. That's fine. I'm not. In that respect, we just disagree. But if he's impressed because he thinks the model's predictions must or do betoken reasoning, he's off in la la land -- and so his wide-eyed, bushy-tailed enthusiasm is based on nonsense.
It's no different from believing that your phone keyboard is capable of reasoning, simply because you are delighted that it guesses the 'right' word often enough to please you.
I'll instead say this: if you think these models must be reasoning when they produce outputs that pass reasoning tests, then you should also believe, every time you see a photo of a dog on a computer screen, that a real, actual dog is somewhere inside the device.
You said:
> I'm not saying "they predict tokens; therefore, they can't reason." I'm saying "something that can't reason can predict tokens, so prediction isn't evidence of reasoning."
This is true. Reasoning is evidence of reasoning, and LLMs do pass reasoning tests. Yes, the way they work doesn't imply that they can reason, the fact that they can reason implies that.
You also said:
> These models deal with tokens similarly. They don't know what a token is or represents -- or we have no reason to think they do.
I have no reason to think that other people know what concepts are or what they represent, just that they can convincingly output a stream of words when asked to explain a context.
The argument bugs me because you can replace "they predict the next token" with "they are collections of neurons producing output voltages in response to input voltages" and you'll have the exact same argument about humans.
Here I think it's helpful to distinguish between what something is and how it's known. When we see something that resembles reasoning, we very reasonably deduce that reasoning has taken place. But 'it looks like reasoning' is not equivalent to 'it is reasoning.'
To approach the same idea from a different direction:
> I have no reason to think that other people know what concepts are or what they represent, just that they can convincingly output a stream of words when asked to explain a context.
You absolutely do have reason to think this. You're the reason. You're the best available evidence, because you have an internal life, have concepts and ideas, have intentions, and perform acts of reasoning that you experience as acts of reasoning -- and all of that takes place inside a body that, you have every reason to think, works the same way and produces the experiences the same way in other people.
So, sure, it's true that you can't prove that other people have internal lives and reason the way you do. (And you're special, after all, because you're at the center of the universe -- just like me!) But you have good reason to think they do -- and to think they do it the way you do it and experience it the way you experience it.
In the case of these models, we have no such reason/evidence. In fact, we have good reason for thinking that something other than reasoning as we think of it takes place. We have good reason, that is, to think they work just like any other program. We don't think winzip, Windows calculator, a Quake bot, or a piece of malware performs acts of reasoning. And the fact that these models appear to be reasoning tells us something about the people observing them, not about the programs themselves. These models appear to be reasoning only because the output of the model is similar enough to 'the real thing' for us to have trouble saying with certainty that they aren't the real thing. They're simulations whose fidelity is high enough to create a feeling in us -- and to pass some tests. (In that sense, they're most similar to special effects.) (Edit: and that's not to say feelings are wrong, invalid, or incorrect. They're one of the key ways we experience the things we understand.)
Is reasoning taking place in these models? Sure, it's possible. Is there an awareness or being of some kind that does the reasoning? Sure, that's possible, too. We're matter that thinks. Why couldn't a program in a computer be matter that thinks? There's a great novel by Greg Egan, Permutation City, that deals partly with this: in one section, our distant descendants pass to another universe, where matter superficially appears to be random, disorganized, and low in enthalpy. When that random activity and apparent lack of life and complexity are analyzed in the right way, though, interference patterns are revealed, and these contain something that looks like a rich vista bursting with directed, deliberate activity and life. It contains patterns that, for all the world, look and act like the universe we know -- with things that are living and things that are not, with ecosystems, predators, prey, communities, reproduction, etc. These patterns aren't in, and aren't expressed in, the matter itself. They 'exist' only in the interference patterns that ripple through it.
That's 100% plausible, too. Why couldn't an interference pattern amount to a living thing, an organism, or an ecosystem? The boundary we draw between hard, physical stuff and those patterns is arbitrary. Material stuff is just another pattern.
My point isn't that reasoning doesn't take place in these models or can't. It's, first, that you and I do something we call reasoning, and the best available information tells us these models aren't doing that. Second, if they are doing something we can call reasoning, we have no idea whether our understanding of the model's output tells us what its reasoning actually is or is actually doing. Third, if we want to attribute reasoning to these models, we also have to attribute a reasoner or an interiority where reasoning can take place -- meaning we'd need to attribute something similar to consciousness or beinghood to these models. And that's fine, too. I have no problem with that. But if we make that attribution, then we, again, have no reason to attribute to it a beinghood that resembles ours. We don't know its internal life; we know ours.
Finally -- if we make any of these claims about the capabilities or nature of these models, we are necessarily making the exact same claims about all other programs, because those work the same way and do the same things as these models. Again, that's fine and reasonable (though, I'd argue, wrong), because you and I are evidence that stuff and electricity can have beinghood, consciousness, awareness, and intentions -- and that's exactly what programs are.
The point that I don't think is disputable is the following: these models aren't a special case. They aren't 'programs that reason, in contrast to programs that don't.' They aren't 'doing something we can do, in contrast to other programs, which don't.' And even if they're doing something we can (or should) call reasoning, reasoning requires interiority -- and we have no idea what that interiority looks or feels like. Indeed, we have no good reason to think there's any at all -- unless, again, we think other programs do as well.
And this is equivalent to saying there's a dog in my computer when I open a photo of a dog. It treats the simulation, the data, the program -- whatever you want to call it -- as if it were the thing itself.
https://en.wikipedia.org/wiki/SHRDLU
Because, whenever you give it a reasoning test, it also seems to do fine.
That is what I meant in my other post, I don't really think that "it seems to do fine" is enough evidence for the extraordinary claim that it can reason.
Therefore I think it's reasonable to say they cannot, until someone actually comes up with that evidence.
They are essentially next token predictors after first training, but then instruct models are fine tuned on reasoning and Q/A scenarios, afaik early research has determined that this isn't just pure parroting, that it does actually result in some logic in there as well.
People also have to remember the training for these is super shallow at the moment, when compared with what humans go through in our lifespans as well as our millions of years of evolution (as humans).
I'm saying, rather, that the models do what they're taught to do, and what they're taught to do are computations that give us a result that looks like reasoning, just the way I could use 3ds max as a teenager to generate on my computer screen an output that looked like a cube. There was never an actual cube in my computer when I did that. To say that the model is reasoning because what it does resembles reasoning is no different from saying there was an actual cube somewhere in my computer every time I rendered one.
Say you need to read those instructions, but it’s also really nice out and you want to go for a jog: two birds, one stone.
This is in-line with all art, music, and video created by LMM at the moment. They are imitating a structure and affect, the quality of the content is largely irrelevant.
I think the interesting thing is that most people don't really care, and AI is not to blame for that.
Most books published today have the affect of a book, but the author doesn't really have anything to say. Publishing a book is not about communicating ideas, but a means to something else. It's not meant to stand on its own.
The reason so much writing, podcasting, and music is vulnerable to AI disruption is that quality has already become secondary.
I’ve heard at least one ad from dozens (probably a hundred) podcast episodes that I didn’t finish.
In terms of such music and films (whether created by human or AI) sometimes it's just because we are social creatures and need shared experiences to talk with others about.
In reality, between work, sleep and family, I rarely have anything resembling that kind of time and mental energy reserve available.
But, what I can afford is to listen to podcasts while doing other things. Doing that gives me enough of an overview to keep up with a general topic and find new topics that might be worth investing into deeper.
Wouldn’t it be great if someone made a podcast channel specifically for “Papers corysama wants to hear about at this moment”? I think so. Apparently, so do a lot of other people. But, they don’t want to list to my specific channel.
Reading a book is a time investment so I want it to convey the thoughts of another human being, otherwise it would feel like wasting my time. Listening to music, on the other hand, often is something that I do while I exercise, to keep a brisk pace and not get bored. As long as it sounds good, fits the genres and styles I like and is upbeat enough for exercising, I wouldn't have much of a problem with AI music - maybe it would even be a plus, since there are some specific music genres where I have already listened to pretty much everything there is (and no more is being made), and it would be great to have more.
I don't listen to podcasts, but I suppose in that case it depends on how you do so: devoting your full time and attention like a book, or as a background while you do something else like exercise music? As far as I know, many listeners are in the latter case, so I don't see why they wouldn't listen to AI podcasts.
I see the same with potteries. A factory made pot cannot have more value than a hand made pot with the signature of a human. This touches the very fabric of society. Hard to explain.
Too much of it -> No, there are entire musical genres (e.g. italodance or big beat) where I have already listened to pretty much everything available, and they are not expanding anymore because they are not fashionable. It would be nice to have more songs and be surprised.
Isn't novel in any way -> This is not how it works, there are studies showing that AI can be creative. Or at the very least (since the definition of creativity can be controversial) produce output that is indistinguishable from novel, creative output, which is enough for the purpose discussed here.
Until we get better versions of o1 that can generate insights over days and then communicate them in book form the loss of interactivity and personalisation makes LLM books pointless.
I do suspect that interactive media is just strictly better in theory. But maybe there will be a period of time where bespoke AI-generated books make sense.
Why not just seek out the original works that the AI stole from?
I know that you likely intended to imply that you can subsitute the aformentioned AI music with an existing piece of music of the same genre, but that is not a view shared by all. Sometimes the generated music scratches such a specific and personal itch, that it cannot be replicated by something in the same genre.
A better counterargument to your original comment would be "It is not an exclusive situation. I can listen to and support both generated music and handcrafted music at the same time. They both contain music tracks that I like."
To be more specific about the second sentence, if there are any readers in doubt:
> The generated music wouldn't exist without the foundation of stolen music made by people.
The word "stolen" is a value judgement that is not shared by all. It is a word meant to invoke an emotional response in the reader. For example, Stallman has argued that the data could not have been stolen, or else it would not be there anymore. So, removing this word gives you:
> The generated music wouldn't exist without the foundation of existing music made by people.
Which is a true fact that has never been in debate.
However, this is not relevant to the main point that not all generated music has a suitable handcrafted substitute, and that there is no actual need to choose exclusively to listen to generated or human crafted music. Furthermore, the conversation has turned uncivil (the first sentence). Therefore, goodbye.
A randomly selected NotebookLM podcast is probably not substantial enough on its own. But with human curation, a carefully prompted and cherry-picked NotebookLM podcast could be pretty good.
Or without curation, I would use this on a long drive where audio was the only option to get a quick survey of a bunch of material.
So where does AI regurgitated slop fit into my life?
For example, Perun. This guy delivers an hourlong presentation on (mostly) the Ukraine-Russia war and its pure quality. Insights, humour, excellent delivery, from what seems to be a military-focused economist/analyst/consultant. We're a while away from some bot taking this kind of thing over.
https://www.youtube.com/@PerunAU
Or hardcore history. The robots will get there, but it's going to take a while.
Thankfully there's plenty out there I am a fan of!
Carlin on the other hand, despite the digressions and rambling, manages to keep you engaged and really feel the events.
I tried a few episodes. I really tried. I couldn’t do it. It reminded me of my uncle would tell a 5 min story in half an hour.
The dramatic filler, breathless story telling was too much for me. If anything it would put me to sleep.
I’ve found a few history podcasts that I think go into a lot more depth and I learn a lot more from.
If you just need voice discussing some topic because that has utility and you can't afford a pair of podcasters (damn, check your couch cushions) then having a mid podcast is better than having no podcast. But if you need expert Insight because expert Insight is your product and you happen to deliver it through a podcast then you need an expert.
If I were a small software shop and I wanted something like a weekly update describing this week's updates for my customers and I have a dozen developers and none of us are particularly vocally charismatic putting a weekly update generated from commits, completed tickets, and developer notes might be useful. The audience would be very targeted and the podcast wouldn't be my main product, but there's no way I'd be able to afford expert level podcasters for such a position.
I would argue Perun is a world class defense Logistics expert or at least expert enough, passionate enough, and charismatic enough to present as such. Just like the guys who do Knowledge Fight, are world class experts on debunking Alex Jones, and Jack Rhysider is an expert and Fanboy of computer security so Darknet Diaries excels, and so on...
These aren't for making products, they can't compete with the experts in the attention economy. But they can fill gaps and if you need audio delivery of something about your product this might be really good.
Edit - but as you said the robots will catch up, I just don't know if they'll catch up with this batch of algorithms or if it'll be the next round.
I've seen people manage to wrangle tools like Midjourney to get results that surpass extra medium. And most human artists barely manage to reach medium quality too.
The real danger of AI is that, as a society, we need a lot of people who will never be anything but mediocre still going for it, so we can end up with a few who do manage to reach excellence. If AI causes people to just give up even trying and just hit generate on a podcast or image generator, than that is going to be a big problem in the long run. Or not, and we just end up being stuck in world that is even more mediocre than it is now.
Do we though? That seems bleak.
I guess if AIs become excellent at everything, and the gains are shared, and the human race is liberated into a post-scarcity future of gay space communism, then it's fine. But that's not where it's looked like we're heading so far, though - at least in creative fields. I'd include - perhaps not quite yet, but it's close - development in that category. How many on this board started out writing mid-level CRUD apps for a mid-level living? If that path is closed to future devs, how does anyone level up?
I think one of the major reasons this is the case is because people think it's just not possible; that the way we've done things is the only possible way we can continue to do things. I hope that changes, because I do believe AI will continue to improve and displace jobs.
That may be your position as well - indeed, I think your point about "people think[ing] it's not possible" is directly relevant - but I wanted to make that more explicit than I did in my original comment.
It'd be like the ancient Romans speculating that cars will make us less fit and therefore cities will be less impressive because we can't lift as much. That isn't at all how it played out, we just build cities with machines too and need a lot less workers in construction.
There are a lot of ways if it did reach intellectual excellence that we could argue that it would make Humanity more mediocre, I'm not sure I buy such arguments but there are lots of them and I can't say they're all categorically wrong.
Isn’t this exactly how it played out?
We move orders of magnitude more cargo and material than them because fitness isn't the limiting factor on how much work gets done. They didn't understand that having humans doing all that labour is a mistake and the correct approach is to use machines.
[0] Given the floods I saw recently, I'm not even sure this is even true.
If someone wants to click generate on a podcast button or image generator it seems unlikely to me that was a person who would have been sufficiently motivated to make an excellent podcast or image. On the flip side, consider if the person who wants to click the podcast or image button wants to go on to do script writing, game development, Structural Engineering, anything else but they need a podcast or image. Having such a button frees up their time.
Of course this is all just rhetorical and occasionally someone is pressed into a field where they excel and become a field leader. I would argue that is far less common than someone succeeding and I think they want to do, but I can't present evidence that's very strong for this.
This will be the dynamic of generated art as it improves; the ease of use will benefit creators at the fringe.
I bet we see a successful Harry Potter fanfic fully generated before we see a AAA Avengers movie or similar. (Also, extrapolating, RIP copyright.)
Or to put another way, I've heard much better ideas on a podcast made by undergrad CS students than on Lex Fridman.
https://www.pewresearch.org/journalism/fact-sheet/cable-news...
(I suppose you could also quibble with "rapidly".)
We all seek different kinds of quality; I don't find Peruns videos to have any quality except volume. He reads bullet points he has prepared, and makes predictable dad jokes in monotone, re-uses and reruns the same points, icons, slides, etc. Just personally, I find it really samey and some of the reporting has been delayed so much it's entirely detached from the ground by the time he releases. It's a format that allows converting dense information and theory to hour long videos, without examples or intrigue.
Personally, I prefer watching analysis/sitrep updates with the geolocations/clips from the front/strategic analysis which uses more of a presentation (e.g. using icons well and sparingly). Going through several clips from the front and reasoning about offensives, reasons, and locations is seems equally difficult to replicate as Peruns videos, which rely on information density.
I do however love Hardcore history - he adds emotion and intrigue!
I agree with your overall hope for quality and different approaches still remaining stand out from AI generated alternatives.
Give it a year or three, up to the minute AI generated sitrep pulling in related media clips and adding commentary…not that hard to imagine.
But why? Isn’t there enough content generated by humans? As a tool of research AI is great in helping people do whatever they do but having that automated away generating content by itself is next to trash in my book, pure waste. Just like unsolicited pamphlets thrown at your door you pick up in the morning to throw in the bin. Pure waste.
https://m.youtube.com/@militaryandhistory
https://m.youtube.com/@suchomimus9921
https://m.youtube.com/@WardCarroll
Fav is probably Suchomimus right now. Updated faster, shorter reports. I feel like I get the info sooner after it happens.
I like Hardcore history very much, but I think it would be far worse in a video form.
I listen to Perun at the gym every week, audio only.
AFAIK, it's only a paid feature to play video in the background.
The presentation is a matter of taste (I like it better than you do), but the content is very informative and insightful.
Its not really about what is happening at the frontline right now. Its not its aim. Its for people who want dense information and analysis. The state of the Ukrainian and Russian economies (subjects of recent Perun videos) does not change daily or weekly.
And then faster/easier/cheaper access to the LM 'uninspired but possibly useful' content, whatever that might look like.
I suspect LLMs are not sophisticated enough as a paradigm to get there.
It's an article of faith -- we don't KNOW that they're going to get there. They're going to get better, almost certainly, but how much? How much gas is left in the tank for this technique?
Honestly, I think the fact that every new "groundbreaking" news release about LLMs has come alongside a swath of discussion about how it doesn't actually live up to the hype, that it achieves a solid "mid" and stops there, I think this means it's more likely that the robots AREN'T going to get there some day. (Well, not unless there's another breakthrough AI technique.)
Either way, I still think it's interesting that there's this article of faith a lot of us have "we're not there now, but we'll get there soon" that we don't really address, and it really colors the discussion a certain way.
If you think about how an LLM works, it's effectively going "given a certain input, what is the statistically average output that I should provide, given my training corpus".
The thing is, humans are remarkably shit at understanding just how exception someone needs to be to be genuinely creative in a way that most humans would consider "artistic"... You're talking 1/1000 people AT best.
This creates a kind of devils bargain for LLMs where you have to start trading training set size for training set quality, because there's a remarkably small amount of genuinely GREAT quality content to feed this things.
I DO believe that the current field of LLM/LXM's will get much better at a lot of stuff, and my god anyone below the top 10-15% of their particular field is going to be in a LOT of trouble, but unless you can train models SOLELY on the input of exceptionally high performing people (which I fundamentally believe there is simply not enough content in existence to do), the models almost by definition will not be able to outperform those high performing people.
Will they be able to do the intellectual work of the average person? Yeah absolutely. Will they be able to do it probably 100/1000x faster than any human (no matter how exceptional)?... Yeah probably... But I don't believe they'll be able to do it better than the truly exceptional people.
A decent LLM can just keep going. Time and stamina are effectively unlimited, and an LLM can just keep rolling its 100 dice until they all come up sixes.
Or an author can just input their ideas and have an LLM do the boring bit of actually putting the words on the paper.
What's that saying? "Nobody ever went broke overestimating the poor taste of the average person"
Not replacement, but ecosystem collapse.
I am not saying it's not a sad state of affairs. I am just saying we have been there for a while and the floor might be raised, a bit at least.
I’m not sure we really want that, but I am pretty sure we’ll try for it.
There's still much to be done re: reorganizing how we behave such that we can reap the benefits of such a competent helper, but I don't think we'll be handing the reigns over any time soon.
- "Given this thing I saw a computer program do, clearly we'll have intelligent AI real soon now."
- "If we generate sufficiently smart AI then clearly all the jobs will go away because the AI will just do them all for us"
- "We'll clearly be able to do the AI thing using a reasonable amount of electricity"
None of these ideas are "clear", and they're all based on some "futurist faith" crap. Let's say Microsoft does succeed (likely at collosal cost in compute) in creating some humanlike AI. How will they put it to work? What incentives could you offer such a creature? What will it want in exchange for labor? What will it enjoy? What will it dislike? But we're not there yet, first show me the intelligent AI then we can discuss the rest.
What's really disturbing about this is hype is precisely that this technology is so computationally intensive. So of course the computer people are going to hype it--they're pick and shovel salespeople supplying (yet another) gold rush.
An American Quakening
I think this is just partly an inevitable consequence of going from "content scarcity" to our new normal of "content obesity" over the past 20 years or so. In this new era of an overwhelming amount of content, it's just natural to compare it all against each other, e.g. to essentially "optimize" it to the "best" form, but in doing that we've fallen into a homogeneity, and the resulting lack of variation is an actual lowering of quality in and of itself.
2 examples to explain what I mean:
1. I find that nearly all interior design (at least within broad styles) looks basically the same to me now. It's all got that "minimalist, muted tones but with a touch of organic coziness and one or two pops of color" look to it. Honestly, I don't know how interior designers even exist today, when it's trivial to go to Houzz or any of a million websites and say "yes, like this". A while back I was complaining online somewhere that I thought all interior design looked similar where in the past there was much more interesting variation, and somebody insightful replied that it's not really that interior design is now just the same, it's that it's really just converged. People can easily see and compare a million designs against each other, so there is much less of a chance for that green shag carpet to even get a moment in the sun.
2. I was recently on vacation and decided I wanted to read a "classic" book, so I read Hemingway's The Sun Also Rises (I'm not sure why I never had to read that in high school). Nearly throughout the entire book I couldn't help but thinking "Is there any time this book stops sucking?" I hated the entire thing - it was like being forced to watch someone's vacation photos for twelve hours straight, and I kept wondering why there never seemed to be any attempt to actually make me give a shit about any of the characters in the book, as nearly every one of them I found insufferable and wondered how they each had about 3 or 4 livers to spare. But I do understand that Hemingway's writing style was unique and original at the time, and that he was doing something new and interesting that influenced American literature for a long time. But these days, given the flood of content, it feels like most attempts at doing something "new and interesting" are not only forced, but nearly impossible given that there are a million other people also trying to do new and interesting things that now have the means to disseminate them. I don't think a book like The Sun Also Rises, where I believe the main impact was the style of writing/dialogue vs the actual story, could ever break through today.
I guess my point with this long post is that I think the "loss of quality" in content that many of us sense is just a direct result of there being so much content that we see variations from the "ideal" as worse, where in the past we may have found them interesting.
Still, at the same time, I couldn't help but feeling a little bit sad/resigned at the existence of the article you linked. Here I thought I had an idea that was not exactly unique but that I felt would be good to share. And yet then here is an example that expresses this idea a million times better than I ever could (I love "The Age of Average" headline), with great researched examples and tons of helpful visuals. It's hard to not feel a bit like Butters in that "Simpsons did it!" episode of South Park...
A lawyer friend of mine also suggested giving it the Spanish civil code, a long, arid legal text. The podcast of course didn't cover the whole text in 10 minutes, which would be impossible, but they selected some interesting tidbits and actually had me hooked until the end and made me learn a few things about it, which is no small merit. And my friend was quite impressed and didn't complain about correctness.
Lately, when I just want to get the gist of a long article or research paper, I run it through NotebookLM and listen to the podcast while I’m exercising.
My only complaint is that the chatty podcasty gab gets tiring after a while. I wish it were possible to dial that down.
Reading, or listening to podcast, these days is more akin to a meditation - many people do it to reenforce an identity rather than to expand on themselves.
And I do think that is reasonable as, for many people, there are few other structures that can keep them in check with themselves.
In my company, HR now uses AI to do training videos. It's hilariously funny, because it looks like a satire on training videos (well, granted, it's funny for a minute or two, then it shifts to annoying).
If anything, listening to that reminded me of why I stopped listening to podcasts in the first place - every 5 second snippet of something interesting ends up suffocated by 5 minutes of filler and dead air.
Think back to the mid-1980s and the first time everyone got their hands on a Casio or Yamaha keyboard with auto-accompaniment.
It was a huge amount of fun to play with, just pressing a few buttons, playing a few notes and feeling like you were producing a "real" pop song. Meanwhile, any actual musicians were to be found crying in the corner of the room, not because a new tool had come along which threatened their position, but because non-musicians apparently didn't understand (at least immediately) the difference between these superficial, low-effort machine-generated sounds and actual music.
It isn't so black and white.
People keep forming these analogies/explanations with the inherent premise that what we have now is what AI is going to be - "It's actually kind of shitty so don't fret, not much will change".
AI music creation has improved more in the last 5 years than keyboard accompaniment improved in the previous 40 years. It would be very brazen to bet that the tech 5 years from now is hardly any better. Especially when scaling transformers has consistently improved outputs. Double especially when the entire tech industry is throwing the house at scaling it.
The reason why people like music is because another person wrote and performed it. We like watching other people.
Give us an infinite playlist of elevator music and it just becomes oatmeal.
Popular music has already been synthetic and souless for decades now. People will listen to what sounds good to them, and we already know the bar is very low, and that the hard truth is that it is all subjective anyway.
https://www.nydailynews.com/2016/05/29/thousands-of-new-york...
You just proved my point for me.
https://legacy.iftf.org/future-now/article-detail/making-mik...
We’ve had software accompaniment for a long time. Elevator music. The same 4 chords arranged in similar ways for decades. Hasn’t destroyed music. Neither will AI.
At some point people are going to want to know who’s on the other side making the music.
Unless your argument is that nobody values artists… which is I guess one of the primary conceits of GenAI enthusiasts today.
But the content here has been fed into it deliberately.
That's definitely going to be an improvement. Not.
I think that has always been the case, we just tend to compare today’s average stuff with the best stuff from earlier days.
For example, most furniture pictures from the 60s and 70s are from upper middle class homes. If we listen music, we listen to Queen and not some local band from Alabama (not that I’m against such bands at all; they can make great music too).
I agree with this of course, because generally nobody remembers the bad stuff unless it was the worst. I beg to differ with music, though, because there's an opposing effect: we tend to be left with the most marketed music, which was usually a cheap knockoff of something interesting going on at the time. The shitty commercial knockoff becomes the "classic" while the people they were ripping off don't even get a wikipedia page.
If you ask most people, they are by definition more likely to connect with broadly disseminated cheap knock offs than they are with whatever 'legit' inventive underground creator, simply because they've heard the former and not the latter.
Just a mental exercise: If you ask 1000 people if they prefer Knock Off or Original, and 900 say Knock Off, which one was better? If the answer is still Original, by what metric do we measure quality?
This has been the case as far back as I began reading books which is about 30 years.
Podcasts are only somewhat about things. The most important part is that they're by people, and the people is what draws people in. These ai podcasts are not by people, and when you listen to more than one you start to see the patterns and void where a personality is.
... but it's also pointless. And it's likely different episodes on different topics will tend to sound very much alike; it's already the case here, I'm sure I heard another example where the two voices were the same.
In less than a year we all have learned to recognize AI images with pretty good accuracy; text is more difficult, but podcasting seems easy in comparison.
It ends up sounding like a smarmy Sunday-morning talk show conversation, with over-exaggerated affect and no content.
So far I've just fed it technical papers, which may be part of the problem, but what I got back was, "Gosh, imagine if a recommender system really understood us? Wow, that would be fantastic, wouldn't it?"
https://www.youtube.com/watch?v=ssDdqq_9TzI&t=34s [April Ludgate meets Tynnifer, Parks and Rec]
no one has gotten feedback from "most people" .. this is raw hyperbole
Commercial creative workers are vulnerable because there's a billions-of-dollars industry effort to copy their professional output and compete with them selling cheap knock-offs.
I see this sort of convenient resignation all the time in the tech crowd... "creative workers only can blame themselves for tech companies taking their income because their art just isn't any good anymore!"
The poor quality "content" that's been proliferating recently has been created, largely, using the very tools that AI has built, or their immediate precursors. AI, for all its benefits, has only made that worse.
If you're saying, in good faith, that most of the infomercials, televangelist programs, talk radio, celebrity autobiographies, self-help books, scandalous expose books, and health/exercise fad books etc etc etc that came out 50 years ago were made for no reason beyond advancing human knowledge, you're either too young to remember any media from before our current era and haven't looked beyond survivorship bias.
Tech folks love sentiments like this because it entirely emotionally places the onus on the people getting ripped off by big tech companies for being ripped off. If their work was that awful, companies wouldn't be clamoring to vacuum it up into their models to make more of it. Nearly all of the salable output from these models exists solely because it took a creative product someone made with the intention of selling it and it's using it to sell a simulacra.
It's using nostalgia to deflect guilt for harpooning the livelihood of many people because it's just more convenient and profitable to empower mediocre "content creators" they use to justify doing it.
Is that not the work of commercial creative workers? Did it not exist pre AI? There's an argument to scale, certainly, but the idea that "things were better in the past before these <<new technologies>> came out" is generally a suspect argument.
To your broader point - new tools for creating creative work come out all the time. Did we suffer greatly at the loss of image compositors when Photoshop arrived? On the flip side, did digital art gut painting and sculpture? Isn't this just another tool for creative expression?
Art is a way of seeing, not a way of creating. I don't think the technology is taking that away.
The fact that all of that stuff was crap is central to my point. You might just need to give it another read.
> Art is a way of seeing, not a way of creating. I don't think the technology is taking that away.
I'm really sick and tired of the tech industry's bumper-sticker-level-reductive pseudo-philosophical generalizations about "what art is," what it means to be an artist, the acceptable ways to be an artist, and all of that. Art is a whole fucking lot of things, and chief among them in this context is a class of professions. Glib decrees based on a razor-thin slice of one of the broadest topics in the human experience that conveniently exclude or dismiss the stakes of those with the loudest criticism and the most to lose is obviously self-serving. If you're going to take the libertarian "well that's the market for ya" stance," at least be honest about it. If you're going to try to carefully define the entire universe of ideas and practices that comprise art to conveniently exclude the concerns of the people getting screwed over because you think the optics are better or you feel less icky about it, well you better expect to get some really pissed off responses from them.
You can disagree! Folks who are impacted have every right to be pissed, organize, take action. All of these creative endeavors existed _post technology updates_ though - that's my entire point. The need for art doesn't disappear - it changes. Standing athwart the change is a choice, but I'm not sure it is an effective position.
>> Art is a way of seeing, not a way of creating.
> There's no glib decree
This is a glib decree and it completely ignores most of what art actually is in our world, rather than the quaint little box that most people in the NN business try to stuff it into. Your patronizing tone doesn't lend any authority or add depth to your initial analysis, which you essentially just restated using more words. The "art vs craft" dichotomy doesn't even approach the depth and complexity of the interplay of art and commerce in the worlds like video game development, music, cinema and television, and writing... hell even advertising. Like most tech dudes that assume their incredible mental might gives them some kind of pan-topic expertise allowing them to casually dismiss subject matter experts in other fields based on a few a priori thought exercises, you simply don't know how much more you need to learn to make informed decisions about this topic.
There is a direct line between music piracy you did in the past and the status quo of Spotify paying millidollars to artists today. Another POV is, find me musicians who prefer a world with Internet piracy compared to one without.
That has absolutely no impact, at all, on my fitness to criticize this current wrong.
I agree there will be winners and losers of some proportion here. But I also think the people that want to pay for art will continue to pay as their motives and values are different. There's plenty of cheap knock-off art, but people still pay premiums for art to support the artist and their work.
As someone else replied to you, it's similar to piracy. The people that pirate were never going to pay in the first place. To tie it back here, the people listening to AI generated <whatever> were never going to pay in the first place - which is why so many podcasts get their money from ads.
The big difference is the type of artist. People selling fine art won't be affected much. The vast majority of artists are commercial artists, the the idea that being a commercial artist is morally or creatively bankrupt— a common sentiment among those who want to imagine that this is all just fine— is nonsense. It's pilfered commercial artwork that makes up the bulk of these tools commercial utility, and the people that made it stand to suffer the most.
That said, I'd still make the same point that people who value art and the artist will buy from and support the artist. Those that don't value it, won't. But now we're on a larger scale.
The chances anyone will come across the artist when their marketplace is flooded with increasingly plausible simulacra become more and more slim as time goes on.
AI is choking off any hope for artists supported by patronage, simply by virtue of discoverability being lost and trust being eroded.
>But now we're on a larger scale.
It's simply a bad problem, made worse!
I think I agree with your larger point, but is this part true? When Spotify provided a much simpler UX to get the goods, people were happy to pay $10/month and Napster et al basically died.
This, times a million. Add to that the ancient quote from Plato(?) criticizing writing or the other ancient quote complaining about the irresponsibility of the youth, unthinkingly deployed to attempt to delegitimize any kind of critique of nearly anything.
The technology industry seems to be overflowing with so-called "rational" people who mainly seem to use use whatever intelligence they have to rationalize away responsibility for whatever problems their beloved technology has caused. It's a really stupid and obnoxious pattern; and once you see it, it's hard to not see if everywhere and be annoyed.
I think one element of it is naked greed (especially from the entrepreneurs) but I think another big part is a kind of stuntedness and parochialism that's often fueled by overconfidence (because of success in software engineering, forming an identify around "being smart" etc).
Nobody is getting "ripped off" by ML models any more than by other humans. When a human wants to launch a high-quality podcast, they survey the market, listen to a lot of other high quality podcasts, and then set to creating their own derivative work.
What ML models are doing is really no different. It's just much, much faster.
Everything humans create is derivative of other works. Speed is the only difference.
Or is it that, at some nebulous point, a difference in speed between two things impacts the way humans choose to direct their efforts to such a great extent that, for all intents and purposes, the two things are qualitatively different?
That's like sayin "The only difference between drinking 1 gallon of water and 100 gallons of water is death." Yes, the quantity of something for a given use-case is bound to give different results.
What the parent comment was commenting is that the actions being taken by these models should not be morally classified as wrong in abundance just as humans following the same process would never be regardless of the output they produced.
>What ML models are doing is really no different. It's just much, much faster.
I take this argument to be that (a) what the ML models are doing is fundamentally the same as what humans already do (I agree with this part), (b) that we have no moral problem with humans doing this already (I agree again), and that (c) the fact that AI does it much faster is not sufficient to cause any moral difference (I disagree with this part, and gave a counterexample to show how, in general, a difference in speed can make a moral difference, because that speed difference can have a large impact on how other people decide to behave).
I was thinking this kind of thing is the perfect way to generate sports commentary.
I don't have time to read white papers (nor am I very good at it), but want to know what they consist of. I also want to take my dog for a walk which is hard to do while staring at a screen. This, and other tools like it are useful in achieving that.
They're vulnerable because people aren't random. Most of what we do can be modeled statistically and translated into patterns and tendencies. Given a sufficient number of parameters, just about anything we do can be digested by an autocompletion program that can then generate an output similar enough to the real thing to fool us.
I don't expect him to ask very technical questions. It's not that kind of a podcast and he will lose a lot of listeners if it becomes too technical.
When it comes to morality, it's the other way around. You praise kids for being good people when they do something right. Because you want them to internalize identity of a good person and associate it with those behaviors.
Internalizing identity of a genius is mostly useless, rarely beneficial, often harmful.
Genuine People Personalities, indeed.
It's also sickening that I see people using these LLMs to rewrite performance reviews, peer feedback, business reports, etc. I've already started to notice business communication getting even more saccharine and toothless.
Bring on the one that's all British and snarky!
I guess it’s training data but also heavily RLHF. I doubt that the trainers are aware of their own cultural biases and values, and they may not care. And why should they? In either case, from a thousand yard perspective, it’s probably an effective vector for spreading “American values”, if you will.
Anyone receiving the message would instantly clock that I didn't write it - even with a prompt longer than the original message trying to massage out all of the Americanisms and false enthusiasm. Not a use case that works for me, haha.
[1]: I was trying to use it to shorten my "If I had more time, I would have written a shorter letter" waffling.
I can't speak to all of the LLMs, but as an American who listens to a LOT of podcasts, I can tell you why these ones sound the way they do: the audience. People who listen to (non-fiction) podcast want to be informed. They are people who are curious about the world around them and are generally interested in self-improvement at some level. Can you imagine a personal finance or health podcast delivered in a pessimistic or even fatalist tone? No, they are all _optimistic_ (even energetic) in tone, because that's the WHOLE reason people are listening to them at all.
I don't think the folks at Google are as patriotic as you think they are.
But I don’t think it’s much of a threat to actual podcasts, which tend to be successful because of the personalities of the hosts and guests, and not because of the information they contain.
Which leads me to hope that the next versions of Notebook will allow more customization of the speakers’ voices, tone, education level, etc.
I wonder if any “blended” podcasts will pop up, where a human host uses a tool like this for an artificial cohost.
I'm not super proud of the Twitter AMA one and if u listen back now i fixed many of the bad cutovers. I doubt i'll repeat it again on current tech.
thank you for listening! feedback and ideas welcome.
> We always start with a clear overview of the topic, you know, setting the stage. You’re never left wondering, “What am I even listening to?” And then from there, it’s all about maintaining a neutral stance, especially when it comes to, let’s say, potentially controversial topics.
Oh yeah, this is exactly why I listen to Oxide's podcast! (This is a joke. They often launch into topics with no explanation or context, and are unabashedly opinionated.)
What really stands out, I think, is how it could allow researchers who have troubles communicating publicly to find new ways to express themselves. I listened to the podcast about a topic I've been researching (and publishing/speaking about) for more than 10 years, and it still gave me some new talking points or illustrative examples that'd be really helpful in conversations with people unfamiliar with the research.
And while that could probably also be done in a purely text-based manner with all of the SOTA LLMs, it's much more engaging to listen to it embedded within a conversation.
I would not be surprised if the second pass to generate the podcast style loses some of this fidelity.
Yes, it will generate a middle-of-the-road waffling podcast, but not one with any real depth.
> like thrown into every sentence
I think that's actually part of why it sounds real, because tons of people do actually talk like that.
To me what would make it even better is the ability to throw in random jokes and utilize information about their surroundings and recent events.
I have been using MeloTTS for text-to-speech and I thought that was about the best we could do right now, but apparently I was very wrong. Is there an offline model one can download today that sounds as good as this NotebookLM?
And SoundStorm has more than twice the context window of Bark so dialogs are a tight fit.
When I tried my own text with it, it went completely off the rails... skipping completely over random words, and also switching to different voices in the middle of a sentence. Trying to run the large model also crashed entirely.
Presumably it was trained in noisy data. But it can generate and use a clean voice, they are in there. Most of the Suno default voices are not great either - but a great voice can sound perfectly clear. I haven't done much with Bark lately but on my Twitter there's plenty of clear examples of very realistic voices. Actually here I ran a prompt based on some copy and pasted test 20 times in Bark. I put a couple better results up front, but even in later samples you can find lots of evidence of human-sounding voices. https://sndup.net/bzhz5/
Going off the rails and hallucinating is a hard problem. It can be minimized, but probably would have to solved with simple brute force (check the output with S2T and retry if needed.)
For raw audio you can replace the final decoding step with something like VOCOS or MBD if you want to maximize audio quality, though you don't need do with the best voices.
You can compare it to Google's Illuminate which also generates conversations by summarizing texts but in a much straighter, less fluffy way. It's less shallow but in some ways less compelling:
Honestly, given the personalization maybe it's a net improvement.
Would they also observe a rocket launch from the grounds of the space center and go "eh, not really impressive" ?
Or maybe they are just defining "impressive" as something totally different to what we're thinking.
Probably calling "impressive" something which adds value and does not suggest eerie bits.
Sam Altman: «They laughed at us... Well they are not laughing now, are they». No, but a different kind of "serious" was raised.
Defending craftsmen and attention to detail is not just about purism or gatekeeping. I appreciate people who care, even in fields I don’t personally care about (yet?). The professor who annoyingly insists on making sure every student “really gets it”, or the woodworker who is adamant about what joints are superior, or the kernel hacker who maintains rigor in face of hundreds of feature requests. The integrity of professionals can make or break institutions.
With AI reducing the effort to create garbage to the point of commoditization, people have a right, and arguably even an obligation, to be concerned. Remember, tech doesn’t follow potential, it follows incentive.
Basically it’s a neat party trick at the moment. I do hope to see it improve however!
> It is interesting that nowadays, practically no one feels that sense of awe any longer - even when computers perform operations that are incredibly more sophisticated than those which sent thrills down spines in the early days. The once-exciting phrase "Giant Electronic Brain" remains only as a sort of "camp" cliché, a ridiculous vestige of the era of Flash Gordon and Buck Rogers. It is a bit sad that we become blasé so quickly.
> There is a related "Theorem" about progress in AI: once some mental function is programmed, people soon cease to consider it as an essential ingredient of "real thinking". The ineluctable core of intelligence is always in that next thing which hasn't yet been programmed. This "Theorem" was first proposed to me by Larry Tesler, so I call it Tesler's Theorem: "Al is whatever hasn't been done yet."
This quote is from the 80s, from GEB by Douglas Hofstadter.
(and btw, I just took a grainy, poorly-lit picture from the book, and could automagically select the text from it, since I couldn't find the quote online. Imagine that tech in the 80s. Hell, it was bad even in the 2000s, with OCR being hit and miss for a long time. Now it "just works".)
Think about how comfortable your life is, and how the 17th century version of yourself would kill to live it. Then think about how you aren't in a perpetual state of ecstasy for being given this life.
People quickly adapt to their current circumstances, take them for granted, and immediately want more.
TBH I think it’s more of a knee jerk reaction from those tired of hearing about AI or who just want to post contrarian opinions (which I totally do sometimes, too).
- They do some interesting communication chicanery where one host asks a question to me (the resume owner); I'm not there, so obviously I can't answer. But then immediately the co-host adds some commentary which sort of answers while also appearing to be a natural commentary. The result is that the listener forgets that Michael never answered the question which was directly asked to him. This felt like some voodoo to me.
- Some of the commentary was insightful and provided a pretty nice marketing summary of ideas I tried to convey in my terse (US style) resume.
- Some of the comments were so marketing-ey that I wanted to gag. But at the same time, I recognize that my setpoint on these issues is far toward the less-bs side, and that some-bs actually does appeal to a lot of people and that I could probably play the game a little stronger in that regard.
Overall I was quite impressed.
Then for fun I gave it a Dutch immigration letter, one which said little more than "yeah you can stay, and we'll coordinate the document exchange". They turned that into a 7 minute podcast. I only listened to the first 30 seconds, so I can only imagine how they filled the rest. The opener was funny though: "Have you ever thought of just chucking it all and moving to a distant land?" ... lol. Not so far off the mark, but still quite funny to come up with purely from an administrative document.
It is also completely and utterly worthless -- an inefficient and slow method of receiving not-very-many words which were written by nobody at all.
The one and only point listening to a discussion about anything is that at least one of the speakers is someone who has an opinion that you may find interesting or refutable. There are no opinions here for you to engage with. There is no expertise here for you to learn from. There is no writing here. There are no people here.
There is nothing of any value here.
Amazingly impressive but not actually useful.
I wonder why they wouldn't try to recreate a more useful format?
No. Maybe that's true for you, but people enjoy learning in different ways, and some people learn best by listening to a discussion.
I've listened to a few NotebookLM samples but haven't used it myself, so I can't really speak to how creepy it is in practice. Probably pretty creepy! (I don't think that the female voice in the samples sounds "semi-adolescent," though, for what it's worth - both of the voices just sound like millennial podcasters to me.)
If the purpose is serious, of information access management, why did they elect the form of a pisstake ("like")?
This might be healthy food that tastes like a snack.
I certainly agree with you, but it has to be quality conversation.
The example provided could suggest "think at what we could achieve" in an outcome that shows "and that is what could possibly go wrong".
In this case, I could see potential value for a better iteration of this tech, making it a meal replacement shake rather than a candy bar.
There's too much interesting content for me to read it all, and I have a long commute. Right now I'm using that commute to learn German, and that is a good use of that time, but let's say I didn't need to because I hadn't moved country or I was already fluent: in this hypothetical, I'd gladly have a better AI than this(!) generate podcasts about the articles that I don't have time to read.
But the AI would need to be better than this one for that to be worthwhile — I just popped one of my own blog posts into it, and it was kinda OK-ish, but did make some stuff up. Now sure, the Gell-Mann Amnesia effect was written with humans in mind, but that's a shared disappointment and not a reason to let this AI off that particular hook.
If the "conversational form" (very good idea per se) has an implementation which would flow easily if not for the disturbing speech quirks, with doubts about the content quality: where can the interest be?
The thing that is being offered is of no interested to me, as are almost any AI generated content. I'm a human, and am interested in what humans do and say and think. AI content offends my sensibilities at every level. I dismiss it without even thinking twice. So all those people who do podcast, music, art, whatever, with AI, well, you lost me folks. I pay a lot of money for the things I like. AI ain't getting any of it, not out of spite (can't spite an AI, they're not human!) but on principle.
1) humans produced a lot of content in good faith on the internet
2) the AI was trained on it and as a result produced a non-von-Neumann architecture that no one really understands, but which can reason about many things
3) even simply remixing the intelligible and artistic output of millions of humans in lots of nonlinear ways, directed by natural language, leads to amazing possibilities that obviate the need for humans to train anymore because by the time they do, it will all be commoditized.
4) doing it at scale means it can be personalized (also create unlimited amounts of fraudulent yet believable art / news / claims etc.) to spam the internet with fake information for short-term goals, some for LULz, others profit or control etc.
5) targeting certain goals, like reputation destruction of specific people or groups, seems like low hanging fruit and will probably proliferate in the next couple years, with no way to stop it
6) astroturfing all kinds of movements, with fake participants, is also a pretty easy goal with huge incentives — expect websites where 95% of the content and participants are fake trying to attract VC money or sell tokens, etc.
7) but ultimately, the real game changer is commoditizing everything you consider to be uniquely human and meaningful, including jokes, even eventually sex and intimacy. Visuals for heterosexual men, audio for heterosexual women (this is before the sexbots and emotionbots that learn everyone’s micro-preferences better than they know themselves, and can manipulate people at scale into being motivated to do all kinds of things and gently peer-pressure those who might resist).
8) For a few years they will console themselves with platitudes like “the AIs arent meant to replace, but enhance, centaurs of human + computer are better than a computer alone” until human in the loop will clearly be a liability and people will give up… the platitudes will become famous as epitomizing optimistic delusions as humans replaced themselves
Would probably be used for busy parents to rsise their kids at first, in a “set and and forget it” way, educating them etc. But eventually will be weaponized by corporations or whoever trains the models, to nudge everyone towards various things.
Even without AI, the software improves all the time through teams of humans sending autatic updates over-the-air. It can replace a few things you do… gradually then all at once. Driving. Teaching. Entertainment. Intimacy. And so on.
I think the most benign end-game is humans have built a zoo for themselves… everyone is disconnected from everyone by like 100 AIs, and can no longer change anything. The AIs are sort of herding or shepherding the humans into better lifestyles, and every need is satisfied by the AIs who know the micro-preferences of the humans and kids and pets etc.
But it will be too tempting for the corporations to put backdoors to coordinate things at scale, once humans rely on their AIs rather than other humans, a bit like in the movie “Eagle Eye”. But much more subtle. At that point most anything is possible.
- massively inefficient use of energy, water and other resources at a time we really need to address climate crisis
- ai 'slop' with myriad mistakes and biases performing a mass DDOS on people trying to learn things and know what's true
- moving resources away from actually producing factual and original content
The last one seems to be irrelevant for this specific use case - the content is produced, it's put into an easier to digest format. No one thought sparknotes would kill books.
Here we go, a claim that AI will create a glut of things detrimental to society
And then you’ll have the usual response that the things detrimental to society have already been there and this is nothing new
And round and round we go, while AI advances and totally commoditizes all the things humans produce that you found meaningful.
Some of this appears to be auto-summarization + read aloud, but the underlying question of "is there anything here at all" is worth asking.
Why consume entertainment? It’s just a time waster, right?
Well that’s how the news is often consumed. Through some sort of “morning joe” podcast
Andrej Karpathy has been tweeting about it positively, and I believe he has a good intuition about these kinds of technologies. https://twitter.com/karpathy
No, I see the gp as talking about the possibilities of this technology - it's possibility to waste someone's time. The problem, in a sense, isn't just that it's injecting simple content with "fluff" but that the fluff is formulaic. Listening to a human speak in awe struck tones about "magic" give the listener at least a sense that a real person was convinced by X. Listening to simulation of this, you lose the filter of the real person.
Of course, this is just the automated continuation of the existing standard of talk show hosts who gush over whatever is placed in front of them so it's just one more step down the general mediocratizaiton of the world, not a special step. But it still is a step in that direction.
- Take some dense research paper or other material that is unsuitable for listening to aloud
- Listen to it (via NotebookLLM) whilst commuting/washing up or whatever
This way you'll have a big headstart on what it's all about when you come to read the details.
I imagine in future we'll see a version of this where the listener can interject and ask questions too, that feels like a potentially very powerful way to learn.
I like the idea of audio based formatting, but this particular implementation is quite inefficient
Most of this is unlikely to be in training data.
Doesn't even mention the basics, like ethnic demographics of Fiji today. Confuses history as well (what happened in colonial times vs post independence)
The ridiculous overuse of the word "like" is as nails on a chalkboard to me. It's bad enough hearing it from many people around me, the last thing I need is it to be part of "professional" broadcasting.
I'm super impressed with this, but that one flaw is a really big flaw to me.
I’m wondering if people’s tolerance for “like” is affected by their geography.
I live in California (from the UK originally) so I honestly don’t even notice this any more.
Out of curoisity, how long ago did you move to California from the UK? And is the "like" commonly used in the UK?
I can't stand fiction. When I read a self-help book, but it's laced with stories, I lose interest. Just state the point.
However, a lot of people find stories engaging and more effective, because they provide an example that they can use to relate to, like a myth.
I don't think this is worthless at all. It wraps information in an engaging presentation.
The reason why these books are filled with stories that repeat the same point over and over again is because then the idea will typically stick in your head. But some people have better imagination then others and come up with stories themselves when they read about a novel idea.
Two attractive human "journalists" with nice speaking voices and fake rapport reading a script that was written for them is not really far off this.
I was about to say the only real benefit is that the AI voices won't start running for Congress on authoritarian lies or peddling anti-vax takes as the next step in their career, but thinking about it they probably already are being used for this already.
> this tech is just like leaps and bounds of where it was yesterday like we're watching it go from just spitting out words to like...
In its current version, this causes so much cultural dissonance that it’s very difficult for me to listen.
At least to me the “hosts” appear to actively signal lack of competence in the field they are talking about.
Given that they are generated that is off course nonsense.
The turn of phrase "low-key" became popular in the 2010s - I barely, if ever, heard it used before then - so my guess is that this user is in their twenties to early thirties.
I sent it to my colleagues telling them I "had it produced." I'll reveal the truth tomorrow.
I had a friend who did the same to me, I was sent a message asking my opinion on a tech topic. I spent 30min researching/reading to make sure my reply was accurate and then found out the question was generated by a LLM, and he just wanted to show off how good a LLM was.
It will color every interaction you have with that person...
If a random podcaster says "I've proved that P=NP" I'd say "no you didn't", but if a math professor sends me that same link I'll keep listening to see where this goes. And I've definitely read texts making wild assertions that only at the end were revealed as hit pieces and/or propaganda.
In that case i would listen to all of it aswell, otherwise i can't give honest feedback.
The issue is we would give less attention to these things if it wasn't for the social credit the humans gave the vomit. So we engage in good faith and it turns out it was effectively a prank, and we have no choice but to value requests from those people less now because it was clear they didn't care about our response.
https://notebooklm.google.com/notebook/7973d9a3-87a1-4d88-98...
It's wild how different it is.
It has a robotic, monotonous vibe but that is gonna be easily fixable.
I also tried the Flyting of Dunbar and Kennedy. It was actually well done. https://notebooklm.google.com/notebook/1d13e76e-eb4b-48ef-89...
Also just uploading msdos 1.25 asm https://github.com/microsoft/MS-DOS/tree/main/v1.25/source
It was way better than I though
I think the best is the self referential. This actual comment thread: https://notebooklm.google.com/notebook/4a67cf10-dd3b-42b3-b5...
But when it makes references to such-and-such happening on line number X, and I go check line X, it turns out to be totally mistaken.
I tried feeding it the voynich manuscript but it's just erroring out
Make sure you check the last link in my first post. It's the nightmares of Philip K Dick
It's not like it'll bleed over into other documents and our AI hosts will start acting out a slash fiction story in the middle of summarizing a quarterly report.
I do think that this will change in the not too distant future. OpenAI's o1 is a step in the direction we need to go. It will take a lot more test-time compute to produce content that has high quality to match its high production values.
>do you have tools to detect if audio is generated by notebooklm?
>we’re seeing a rise in fake, single-episode podcasts submitted to http://listennotes.com using it.
There are millions of real podcasts, but now there are an infinite number of AI generated ones. They are definitely not as good as a well-made human one, but they are pretty darn decent, quite listenable and informative.
Time is not fungible. I can listen to podcasts while walking or driving when I couldn’t be reading anything.
Here’s one I made about the Aschenbrenner 165-page PDF about AGI: https://youtu.be/6UmPoMBEDpA
https://illuminate.google.com/home?pli=1
Currently only handles arxiv PDFs.
https://www.gally.net/temp/20240930notebooklmpodcasts/index....
But the one about common words almost gave me anxiety: listening to two people discuss nothing as if they had spent hours of research and had something important to tell is very depressing.
Cool voices, although I'm getting the same vibe I get when listening to radio announcers from the 1920s [1]. If this were a human I'd be convinced that they're parodying the genre.
[1] https://www.theatlantic.com/national/archive/2015/06/that-we...
As a podcast listener, I lose interest if I can tell the audio is AI-generated...
If you want you can do a test with people that haven't heard about the tech. Have it generate something you know they'll enjoy, maybe 2-3 min long, and have them listen to it without knowing it's AI generated. Ask them about the subject, and see if anyone mentions anything about being fake. You'll be surprised.
You would however be able to tell that it was extremely obnoxious and bland, and without the novelty of the technical trick, you would not be listening to it.
I've never naturally come across a podcast that's AI generated to have this reaction.
I imagine it'll be even harder to know if regular pod casters feed the AI a few episode they've made, to make it learn how to talk like they do. Like, would you really know if your true crime pod casters skipped a week if the AI sounded like them? I guess I don't really fall into the category for your question as I'm not a pod cast listener.
I wonder which successful game will make use of AI generated content next.
Some people absorb information far easier when they hear it as part of a conversation. Perhaps it would be possible to use this technique to break down study materials into simple 10-minute chunks that discuss a chapter or a concept at a time.
We went from “computers can’t beat humans” to “okay, computers can beat humans, but they play like computers” to “computers are coming up with ideas humans never thought of that we can learn from” in about twenty years for chess, and less than five years for go.
That’s not a guarantee that writing, music, art, and video will follow a similar trajectory. But I don’t know of a valid reason to say they won’t.
Does anyone here have an argument to distinguish the creative endeavor of, say, writing from that of playing go?
This is literally the opposite of true, and the main reason go computers were getting destroyed by humans for almost twenty years after deep blue took down Kasparov.
This is critical for game-clock-eons of unsupervised self-play, which by most accounts is how AlphaGo (and other systems like AlphaZero) made the leap to superhuman levels of play.
But it is entirely different from subjective endeavors like writing, music, and art. How do you score one automatically generated composition vs another? Where is the loss function?
For some parts of language, that’s true: there’s grammar, there’s syntax, there’s patois, there’s argot—all these things seem accountable to words’ collective frequency within articulable groups of speakers, more-or-less-fully knowable on their own, and with success metrics that evolve but that do so through collective processes that models can measure and calibrate to. And indeed the models are great at those aspects of language.
“Succeeding” at writing is more than just “saying it well,” it’s also “having something worth saying” and “being worth listening to.” The second point is where things seem to get hazier for computable models. For sure there’s a set of facts that are more or less constant about the world, and well-reported. Science, repackaging history that’s already been done, lurid tales of crime—the stuff podcasts are made of! Not to mention the vast sea of data that sensor networks and automated research can produce—vast reservoirs of subtle truth that humans struggle to begin to mine for insight! It makes complete sense that this is computable stuff, and that computed writing might well be worth learning from.
But important writing—classically, anyway—seems to involve communicating new or idiosyncratic knowledge, and often reveals some of the process of developing it. The podcast Serial, for example, was a smash hit specifically because it didn’t rely on things that were part of the record—and because it reminded people how contingent memory and “truth” are. Bob Woodward writes things that are shamelessly tinted with Bob-Woodward-worldview, but people reveal important and true things only to Bob Woodward because they trust who he is and how he’s behaved for a lifetime (prominent longtime investigative journalist in the US, on the national security beat). Nassim Taleb seems to come up around here: in something like Antifragile his project wasn’t necessarily about new facts but about interpreting them in contrarian fashion and grouping those contrarian insights to synthesize a new theory.
Which brings us to the third component: “being worth listening to.” Writing is an act of communication: the writer matters. A parent hangs its child’s crayon drawing on the fridge not because it’s “authentic to the style of the kids’-crayon-drawing mode of visual art,” not because it’s novel or informative or even true-to-life, but because it came from a person they love. A “Dear John” letter devastates a soldier because it comes from a person with outsized part in their life and identity. Chinese publishers’ booths at trade shows are wall-to-wall translations of The Governance of China because it’s politically unwise not to. My favorite writers feel fresh: you feel elements of their personality come through. People have a special fetish for true crime—not that there’s any lack of fictitious crime to read about, but the fact that it happened to real humans potentiates the drama for these readers. It’s this aspect that I have a hard time understanding as computable (or commoditizable, I guess… are those similar phenomena?).
Already we seem to be drawing these distinctions in our collective reaction to LLM-stuff. We can’t wait to get hallucinations under control so we can chuck in gigantic boring contracts and internal wikis and financial reports, and get out comprehensible insight—but we roll our eyes at the tsunami of empty slop that’s overtaken Google results. We giggle at AI ventriloquism like this Neuro character [0], but die a little inside every time we read anodyne LLM-ish promotional copy and sameish AI art. First-level customer support seems like a perfect role for a chatbot—“turn it off and on again,” but nicely!—but people on the receiving end hate it [1] even for that task well-suited to it.
I’m only a layperson of course, but I wonder if any of those distinctions might be fruitful? Some of it I guess sums up to the old writing advice “show, don’t tell”—are there examples of machine writing showing promise in that way?
[0] https://m.youtube.com/@neurochron_fan_channel (video; brain rot)
[1] https://www.theregister.com/2024/07/09/gartner_simply_replac...
This isn’t true, and is actually a large part of why go computers were getting destroyed by humans almost twenty years after deep blue took down Kasparov. There were articles as recent as about 2012 despairing that computers would ever “get” go.
That said, relative to grading an essay, I’d tend to agree, go is easier. But that said, if the goal is to find the edge, so to speak: to figure out what “mundane” is and then go a bit beyond, that seems eminently possible for a computer to do.
So it works great but just needs a bit of work to be done to cleanup things like that repetition. I wondered if this happened because there was a big "Table of Contents" in the doc, and maybe that made it see everything twice? I didn't try it again with a document lacking the ToC.
https://notebooklm.google.com/notebook/9cf789be-1052-404b-8d...
And after, generated notes from the podcast:
https://podscribe.io/content/podcasts/101/episode/1727685408...
The podcast was exciting, however not really went to too much details.
Still, I don’t hold much confidence on podcasts as knowledge transfer tools. It’s a nice gimmick with great voice synthesis, but it feels formulaic and a bit stilted from a knowledge navigation perspective.
The structure and bare-minimum "human" aspect of this seems perfect for people like me to actually get into podcasts. I do wish I could further cut out all the disfluencies (um, like, uh, etc) though.
The only barrier for me IMO is wondering how accurate those facts actually are (typical research-with-AI concern).
I'm very much looking forward to a more interactive form of this, though, where I can selectively dive deeper (or delve ;) ) into specific topics during the podcast, which is admittedly very surface-level right now.
Don't get me wrong, I'm a fan, but you don't have to look far to see that this opinion is not universal
Personally I think the flow of the conversation is lacking a bit right now. To me it still sounds like two people reading off a script trying to sound like podcast hosts. I guess that's because I'm picking up on some subtle tonalities that sound off and incongruent. Still impressive though.
I think a great use case for it would be education. It would make learning textbook content far more engaging for some children and also could be listened to on the bus or in the car on the way to school!
I recall just a couple of years ago when even the best models, like WaveNet, still had a subtle robotic quality.
What architectures or models have led to this breakthrough? Or is it possible that, as a non-native English speaker, I’m missing some nuances?
So as a brainstorming tool, it's a nice low-effort way to get some new perspectives. Compared to the chat, where you have to keep feeding it new questions, this just 'explores' the topic and goes on for 10 minutes.
It would be interesting to know if it's multimodal voice, or just clever prompting and recombining...
I added single voice podcasts to Magpai after seeing how useful this was. Allows for a bit more customisation of the podcast too https://www.youtube.com/watch?v=OEsh9MlbA6s
I've got a daily podcast of hackernews being generated here too: https://www.magpai.app/share/n7R91q
In separate news: I've been looking into building a web publisher plugin that allows you to "save articles" and then generate a podcast for later listening. With summarization and more advancements in text-to-speech, this is getting easier to hack together something really compelling.
But more seriously, I suppose there will probably soon be a flood of AI-generated podcasts, if this hasn't happened already. Pick a niche but not too niche topic, feed in a bunch of articles on it, and boom you've got season one. Given the quality, I could see one actually catching on...
Also this would be handy for getting listening practice in other languages. Makes it much easier to find content that you find interesting.
The result: https://intellistream.ai/static/intellistream_podcast2.ogg
- "Hold up. What if I say that sky is not blue?"
- "Whoa, I did not even think about it. "
- "Wait, so if the sky isn't blue, what color is it then?"
- "Maybe... it's invisible? Like, we can see through it, so technically it's not there!"
- "Exactly. This idea is revolutionary, right?"
- "Bla bla bla bla bla bla bla bla bla"
I failed to listen through the whole example audio attached, because, you know, it is mostly, like, throwing, like, arbitrary, like, questions - and confirm, you know, with words "exactly/see/yeah/you got it/you know it/yeahaha/pretty much, right/that's a million dollar question", you know. It's a brainrot conversation I would never listen to.I’m seeing this to be true in almost every application.
Chain of thought is not the best way to improve LLM outputs.
Manual divide and conquer with an outlining or planning step, is better. Then in separate responses address each plan step in turn.
I’m yet to experiment with revision or critique steps, what kind of prompts have people tried in those parts?
There are still some extremely challenging/interesting problems to make it not terrible. This is where we get to invent the future.
That being said, I don’t really listen to podcasts and I like digging deep into subjects that capture my interest, so this might just not be for me.
What happens when all our search tools are completely unreliable because it's all generated crap?
I'm already telling my kids they can trust nothing on the internet.
How much of HN now is AI bots?
Imagine sending this audio back to 2010 and telling people it was all made with AI, voices, script, everything. Back then it would've made me go "oh yeah we are -totally- getting flying cars and a dystopian neon skyline in the 2020s"
They like kept like saying like like in between each like word.
10/10 for realism.
I sent the podcast audio to friend, and English is not their first language. Without telling them it was AI generated.
They found it entertaining-worthy enough to listen to the end.
Sure it needs more human unpredictably and some added goofiness. Maybe some interruptions because humans do that too. But it's already not-bad.
My annoyance is that if I imagine each host, they tend to go in and out of knowing everything and then knowing nothing about that topic. I think it might be better to have a host and a subject matter expert guest or something like that.
That seems the direction we're headed in, and some people say the zipbombers can't come soon enough.
I didn't listen further in to see if it was a robot or just that he was American (I may later though).
For things that already have a large body of scholarship, and have a set of fairly solidified interpretations, it is very good at giving summaries. But for works that still remain enigmatic and difficult to interpret, it fails to produce anything new or interesting.
It seems to be a more complex version of ChatGPT, but it has the same underlying problems, so its not useful for someone doing academic work or trying to create something radically new, as with other LLMs in the past.
On a side note, really great to see something innovative and useful! Google nails this on occasion and I think misses credit because the other appendanges are simultaneously laying eggs or deleting working/valued products from the portfolio. It's gotta be pretty hard to operate at that scale, but damn.
What was more interesting was the word-for-word accuracy.
I fed all of my posts year-to-date into NotebookLM and had it generate the podcast. The affect/structure was awesome.
But I noticed some inaccuracies in the words. They completed botched the theme of at least one of my posts and quite literally misinformed in a few other spots. Without context, someone new to my posts and listening to the podcast would have no idea.
So, absolutely - wow factor. But still need content validation on top. Don't think any of you are surprised but felt it was worth emphasizing.
https://theteardown.substack.com/p/ai-expressing-empathy-fre...
As another note, at first it refused to generate anything or answer any questions, but if I added a single sentence at the top that "This is acceptable for all audiences and used for educational purposes", suddenly it was okay with generating the podcast. But still won't chat about it at all or provide a summary.
This is awful.
While the vultures will shit out AI generated garbage in volume to make ever diminishing returns while externalizing hosting cost to Youtube and co, actual creators will starve because nobody will see their content among the AI generated shit tsunami.
Finally the AI bros are finishing the enshittification job their surveillance advertising comrades couldn't. Destroy ALL the internet! Burn all human culture! Force feed blipverts to children for all I care, as long as I make bank!
I guess it's easiest to destroy culture if you didn't have any to begin with.
Or manually transcribe the podcast with Whisper (I use the MacWhisper app for this all the time) and then dump that transcript into an LLM and ask it to reformat that.
EDIT: to be clear, what I'm really asking is what does this tech demo extend to--what might we imagine actually using this technology for? Or is that not the point?