How else do you think the "up next" and "recommended videos for you" feature works? Did you think they were serving up relevant videos purely out of chance? If anything a dumb rule like "highlight the subscribe button if the transcript mentions it" is much more benign then a recommendation engine that can somehow always serve you a relevant video.
It wouldn't surprise me if they use transcripts, but neither feature strictly requires it.
I think the real risk here is users extrapolating from "YouTube can automatically respond to the content of videos". The Mr Beast thing is still blowing up, could they automatically identify illegal lotteries? Could they identify things that aren't suitable for YouTube Kids but end up there anyways? So on and so forth, YouTube's moderation is basically always under fire.
This seems like a risky mixed message to send. Being able to light up the Subscribe button doesn't imply the technical ability to do better moderation, but I also wouldn't want to be the guy that has to explain that and how it works to the Senate.
Don’t get started with the “it’s a slippery slope” kind of shit. Short of living in some society where it’s every man for himself, kids shouldn’t be playing lotteries
It also helps that there’s no counterparty risk with the like button. I don’t think anyone is trying to “trick” that feature, and it wouldn’t matter if they did.
I sure wish they did.
Understanding the content is ultimately the only way to properly enforce and discriminate IP theft vs. fair use at scale. And the Senate wants that feature as the IP lobby is strong.
All the different mechanisms that youtube has used is an interesting rabbithole but its ultimately an unwinnable cat and mouse game unless you have an algo that understands the actual content.
Even the human reporting has become weaponized. Don't like the message of a video? have a bunch of youtube accounts flag it and it gets demonetized.
OotL; Mind sharing a link or two?
Familiar with MrB but no idea about anything blowing up. Quick search doesn't make the referenc obvious.
That’s the original video that kicked it off. It’s pretty mundane as far as internet controversy goes. The lotteries were run poorly and as such were probably illegal, not exactly uncommon on YouTube. The show largely targets kids, so there’s some accusations of it effectively being gambling for kids.
After that people dug into his whole life and found out one of his employees was convicted of statutory rape and he was aware, which is a bad look for a channel mostly watched by children.
It’s like a 3/10 on internet controversy scales. Doing something common and dumb but not obviously malicious, and then just stuff that looks bad without concrete evidence anything bad actually happened.
It’s not terribly interesting, I only saw because of a comedian making fun of it.
I never get any relevant ads or videos for me :(.
It is absolutely possible to serve recommendations in a content agnostic way based on other user behavior. That is how the Youtube's recommendations worked at least initially. It should be obvious that they didn't have this ability to analyze content when Youtube started almost 20 years ago.
And even in a content aware system, there is a difference between matching up keywords and understanding what those words mean. That is why an automatic transcript feels more neutral. That just transcribes what is being said. There is no implication of any analysis on the meaning of what is being said.
>If anything a dumb rule like "highlight the subscribe button if the transcript mentions it" is much more benign then a recommendation engine that can somehow always serve you a relevant video.
I'm not saying this feature is scarier than the recommendation engine. I'm saying Google revealing how much they know about the content of the video makes everything they do, including their recommendation engine, a little scarier.
CSS animations werent there I guess, but given a manual transcription, either by the submitter, or by a person watching, youtube definitely had the capability to do exactly this, maybe with the button being some flash animation.
I mean the "given a manual transcription" condition means the very much didn't have "the capability to do exactly this". The animation itself isn't the interesting part here.
Yes but the quality and accuracy of that system would be inferior, with a long delay until it gathers enough collaborative filtering information to be useful.
Indeed, and I assume that's why the home page is full of sports and music when I'm not logged in. (Or was recently, it now shows a blank page).
I'm a very strange person: I don't have any interest in sports, and will listen to perhaps 2 or 3 bits of music every week — if I've heard a piece before, I generally find it predictable and by extension dull, unless some prior association of positive feelings can override that.
1) similarity to other users (people who watched this also watched)
2) similarity in words in title or description (or the generated transcript)
How did you think they did automated closed captioning..?
Think of it like reading a foreign language that uses the same alphabet. You can give me some French text and I can read it well enough for a French speaker to understand it. That doesn't mean that I myself understand what I read. Transcripts are the former and this feature is the latter.
I'd expect the button to light up at random times in a video about "smash the like button" causing the button to light up, rather than in the end call to action for the video where the video author actually wants the watcher to click the button.
Similarly, I doubt it can properly understand when the video presenter implies the term "smash the like button" while instead say, leaving an empty space where they would usually say such a thing.
If I pointed you at some french text and asked you to point at the term "le Chateau" you'd be able to do that without understanding french.
> It's just a substring match.
Let's just say this is true. That is a super simply process, but what would it look like?
Step 1: Transcribe the audio into text
Step 2: Run substring match on text
The transcribing/close captioning feature only does step 1. This shows that a step 2 is possible. I think you would have to be naive to think the capability to do this type of analysis on the transcribed text was designed for only this feature and would never be used for anything else. This feature is announcing that Youtube isn't merely creating transcripts of the audio in videos, it is running some unknown amount of analysis on that data.
As I said in my original comment, "it isn't that I didn't know Google had this ability", but this is literally a glowing sign pointing to this fact. I think the danger of Google reminding people of this has potential to outweigh the benefit of the "that's cute" reaction that this is designed to elicit.
Why would you have to be naive to believe this? The subtitles, with timings, are available on the client-side already. You seem to allude that doing this work would require some sort of deep analysis work. I think it's really more like 5 lines of JS, and 4 of them are producing the fun animated gradient :P
This feature is like walking into your kitchen one day to find a dead cockroach. I’m saying that is an indication of a cockroach problem while you’re effectively responding with “it’s fine, the one cockroach is dead and there is no reason to believe there are any others.”
> This feature suggests Google understands what was said.
Running .includes() on the client does not imply that Google has any "understanding" of what was said. It only implies that they ran .includes() on the client. includes() does not "understand" anything.
The thing I really don't understand is this: the fact that Google has closed captions at all implies they do an enormous more "analysis" than this minor feature could possibly require. If you understand how Google does CCs and what that means, this shouldn't have bothered you at all.
In your analogy, it's like you see a mound of a hundred thousand cockroaches, but you're are worried about a dust speck in another room.
Do you want to have a good faith conversation about this? Because going back to debating the meaning of "understanding" after I already said this was misleading is not a good way to have a conversation.
>The thing I really don't understand is this: the fact that Google has closed captions at all implies they do an enormous more "analysis" than this minor feature could possibly require. If you understand how Google does CCs and what that means, this shouldn't have bothered you at all.
Can we set up a baseline that there is a difference between content agnostic analysis and content aware analysis? Transcripts are content agnostic in that they can be produced without any comprehension of the words said. This feature is content aware in that it is looking for specific meaning in the words said. Do you not see any difference between these two?
Regardless, there is in my opinion a clear distinction in sophistication between a filter and something that triggers a timed action. And that was really what my original comment was about, this feature's elevated sophistication is a conscious reminder of Google's capabilities. Normally that is out of sight and out of mind which is probably better for Google.
Call it "understanding", call it "content aware analysis". I guarantee that their closed captioning service has much more of that quality than this new feature does.
> Can we set up a baseline that there is a difference between content agnostic analysis and content aware analysis? Transcripts are content agnostic in that they can be produced without any comprehension of the words said. This feature is content aware in that it is looking for specific meaning in the words said. Do you not see any difference between these two?
Again, I don't see it. CCs are not content agnostic: they have to have semantic understanding of the words said in order to produce accurate results. How do you think CCs differentiate between the words "to", "too" and "two" without looking at the surrounding words and having some idea of contextual usage? How do you think CCs can tell between "there" and "they're" without understanding if the speaker is referring to a person or a location? This is only the tip of the iceberg as to how CCs actually work, and more "content aware analysis" will always lead to more accurate CCs.
Still can't get away from that "understanding" debate. You're also now equating an understanding of context with an understanding of meaning. An understanding of meaning isn't needed to differentiate between "to", "two", and "too" because they're all used differently in sentences. When the system encounters those, I don't think it goes to the definitions and tries to find which word makes the most meaningful sentence. Most times the specific homophone can be inferred based on things like part of speech and the part of speech can often be inferred from a sentence without knowing any meaning.
For example, would the system be able to properly handle homophones that are grammatically similar? Could it consistently transcribe sentences like "I have Celiac disease and enjoy the taste of rose water, so I prefer flower to flour in my deserts." That is an easy sentence to understand for anyone who knows the meaning of those words, but there are no grammatical or structural indications on which flower/flour to use.
But either way, that is getting way too deep in the weeds compared to where my point started. This feature calls attention to an analysis of meaning because the user sees the software reacting to the meaning of the content of the video. A transcript does not call attention to an analysis of meaning because the behavior of the software does not change based on the content of the video.
Your first comment - the one that started all this - was, as far as I can understand, arguing that this feature indicated that Google had the capabilities to do more advanced - understanding? processing? meaning analysis? - than it had done in the past. If I keep coming back to that, well, it's because it appears to be your main point. If it's not, please correct me.
> Most times the specific homophone can be inferred based on things like part of speech and the part of speech can often be inferred from a sentence without knowing any meaning.
This is not true. I don't think I have enough more responses on HN to fully explain why homophones can not be inferred without understanding meaning, but I encourage you to go and read about how transcription works!
> For example, would the system be able to properly handle homophones that are grammatically similar?
I mean, this is easy enough for you to check. Here's some videos about flour / flower - notice how the CCs correctly determine if the word is flour or flower with almost 100% accuracy.
https://www.youtube.com/watch?v=y8vLjPctrcU https://www.youtube.com/watch?v=xdaRvErv2Kc
> This feature calls attention to an analysis of meaning because the user sees the software reacting to the meaning of the content of the video.
Are you saying you specifically think that YT is analyzing meaning from this feature, or just some generic user? I think you are smart enough to know that it's not true, but perhaps my mom might not understand that CCs require infinitely more processing power and this feature is just a drop in the bucket. (If you really still don't think it's true, definitely go read more about how CCs are made!)
Here is what I said. "It highlights how much Google analyses the content of its videos... It isn't that I didn't know Google had this ability...". My point was not that I learned about Google's capability from this feature or that this capability was new, it is that this calls attention to Google looking for meaning in the content of the video. A transcript does not call attention to Google looking for meaning regardless of how the transcripts are prepared.
>I mean, this is easy enough for you to check. Here's some videos about flour / flower - notice how the CCs correctly determine if the word is flour or flower with almost 100% accuracy.
Both those videos include the correct homophone in the title and description of the video. Choosing the correct one is not an indication of the system using the meaning of those words, it is pattern recognition. Every use of "flower" means the next usage is less likely to be "flour". The specificity of the example I used was important because it used both "flower" and "flour" in a way that can only be distinguished by the meaning of the words.
>Are you saying you specifically think that YT is analyzing meaning from this feature, or just some generic user? I think you are smart enough to know that it's not true, but perhaps my mom might not understand that CCs require infinitely more processing power and this feature is just a drop in the bucket. (If you really still don't think it's true, definitely go read more about how CCs are made!)
This feature is a glowing sign that Youtube as a company analyses the content of the videos for the meaning of what is said in those videos. You are too deep into the technical details trying to assign credit for what aspect of Youtube does the "understanding" or which "require[s] infinitely more processing power".
Think of this feature like receiving mail and you see one of the letters has already been opened. That could make you feel like your privacy was invaded in a way you wouldn't feel after receiving a postcard. And now we have spent several comments debating whether a torn envelope indicates whether anyone read the letter and whether a postcard is private.
Try 1 to 2.
I honestly suspect this feature is just implemented as a side effect as the auto-captioning... and now I'm a bit curious if carefully crafted manual captions can trigger it to go off continuously.
There's a whole bunch of lore about the types of words and phrases (such as slurs and profanity) that can affect algorithm rankings, monetization, or child-restriction.
Doing a regex search for "smash that like button" is trivial by comparison.
So that sweet engagement KPI is positively impacted and Youtube's CRO/experiment/Ab testing team gets their cookies.
There are probably various small tweaks like this in the UI.
The biggest issues tend to be when "Big Brother" style companies perform that kind of analysis on your content like the Microsoft AI screenshotting backlash.
I think I've missed this - what happen?
Google analyses the content of the uploaders' videos yet "YouTube search", from the end user's perspective, is seemingly searching strings in titles, maybe tags. These titles generally suck for this purpose, even containing unicode garbage. Throw in all the "recommended" crap in the "results", videos that do not contain the strings being searched, and it barely feels like search at all. The end user is sent "results" she never asked for.
I mean, it has autogenerated captions, translated (which i use to learn languages, fun)