Yes, it will generate a middle-of-the-road waffling podcast, but not one with any real depth.
Yes, it will generate a middle-of-the-road waffling podcast, but not one with any real depth.
Honestly, given the personalization maybe it's a net improvement.
Would they also observe a rocket launch from the grounds of the space center and go "eh, not really impressive" ?
Or maybe they are just defining "impressive" as something totally different to what we're thinking.
Probably calling "impressive" something which adds value and does not suggest eerie bits.
Sam Altman: «They laughed at us... Well they are not laughing now, are they». No, but a different kind of "serious" was raised.
Defending craftsmen and attention to detail is not just about purism or gatekeeping. I appreciate people who care, even in fields I don’t personally care about (yet?). The professor who annoyingly insists on making sure every student “really gets it”, or the woodworker who is adamant about what joints are superior, or the kernel hacker who maintains rigor in face of hundreds of feature requests. The integrity of professionals can make or break institutions.
With AI reducing the effort to create garbage to the point of commoditization, people have a right, and arguably even an obligation, to be concerned. Remember, tech doesn’t follow potential, it follows incentive.
> like thrown into every sentence
I think that's actually part of why it sounds real, because tons of people do actually talk like that.
To me what would make it even better is the ability to throw in random jokes and utilize information about their surroundings and recent events.
I have been using MeloTTS for text-to-speech and I thought that was about the best we could do right now, but apparently I was very wrong. Is there an offline model one can download today that sounds as good as this NotebookLM?
And SoundStorm has more than twice the context window of Bark so dialogs are a tight fit.
When I tried my own text with it, it went completely off the rails... skipping completely over random words, and also switching to different voices in the middle of a sentence. Trying to run the large model also crashed entirely.
Presumably it was trained in noisy data. But it can generate and use a clean voice, they are in there. Most of the Suno default voices are not great either - but a great voice can sound perfectly clear. I haven't done much with Bark lately but on my Twitter there's plenty of clear examples of very realistic voices. Actually here I ran a prompt based on some copy and pasted test 20 times in Bark. I put a couple better results up front, but even in later samples you can find lots of evidence of human-sounding voices. https://sndup.net/bzhz5/
Going off the rails and hallucinating is a hard problem. It can be minimized, but probably would have to solved with simple brute force (check the output with S2T and retry if needed.)
For raw audio you can replace the final decoding step with something like VOCOS or MBD if you want to maximize audio quality, though you don't need do with the best voices.
Basically it’s a neat party trick at the moment. I do hope to see it improve however!
> It is interesting that nowadays, practically no one feels that sense of awe any longer - even when computers perform operations that are incredibly more sophisticated than those which sent thrills down spines in the early days. The once-exciting phrase "Giant Electronic Brain" remains only as a sort of "camp" cliché, a ridiculous vestige of the era of Flash Gordon and Buck Rogers. It is a bit sad that we become blasé so quickly.
> There is a related "Theorem" about progress in AI: once some mental function is programmed, people soon cease to consider it as an essential ingredient of "real thinking". The ineluctable core of intelligence is always in that next thing which hasn't yet been programmed. This "Theorem" was first proposed to me by Larry Tesler, so I call it Tesler's Theorem: "Al is whatever hasn't been done yet."
This quote is from the 80s, from GEB by Douglas Hofstadter.
(and btw, I just took a grainy, poorly-lit picture from the book, and could automagically select the text from it, since I couldn't find the quote online. Imagine that tech in the 80s. Hell, it was bad even in the 2000s, with OCR being hit and miss for a long time. Now it "just works".)
Think about how comfortable your life is, and how the 17th century version of yourself would kill to live it. Then think about how you aren't in a perpetual state of ecstasy for being given this life.
People quickly adapt to their current circumstances, take them for granted, and immediately want more.
TBH I think it’s more of a knee jerk reaction from those tired of hearing about AI or who just want to post contrarian opinions (which I totally do sometimes, too).
You can compare it to Google's Illuminate which also generates conversations by summarizing texts but in a much straighter, less fluffy way. It's less shallow but in some ways less compelling: