The YouTube like button glows when you say "smash that like button" [video]
youtube.com
youtube.com
How else do you think the "up next" and "recommended videos for you" feature works? Did you think they were serving up relevant videos purely out of chance? If anything a dumb rule like "highlight the subscribe button if the transcript mentions it" is much more benign then a recommendation engine that can somehow always serve you a relevant video.
It wouldn't surprise me if they use transcripts, but neither feature strictly requires it.
I think the real risk here is users extrapolating from "YouTube can automatically respond to the content of videos". The Mr Beast thing is still blowing up, could they automatically identify illegal lotteries? Could they identify things that aren't suitable for YouTube Kids but end up there anyways? So on and so forth, YouTube's moderation is basically always under fire.
This seems like a risky mixed message to send. Being able to light up the Subscribe button doesn't imply the technical ability to do better moderation, but I also wouldn't want to be the guy that has to explain that and how it works to the Senate.
Don’t get started with the “it’s a slippery slope” kind of shit. Short of living in some society where it’s every man for himself, kids shouldn’t be playing lotteries
It also helps that there’s no counterparty risk with the like button. I don’t think anyone is trying to “trick” that feature, and it wouldn’t matter if they did.
I sure wish they did.
Understanding the content is ultimately the only way to properly enforce and discriminate IP theft vs. fair use at scale. And the Senate wants that feature as the IP lobby is strong.
All the different mechanisms that youtube has used is an interesting rabbithole but its ultimately an unwinnable cat and mouse game unless you have an algo that understands the actual content.
Even the human reporting has become weaponized. Don't like the message of a video? have a bunch of youtube accounts flag it and it gets demonetized.
OotL; Mind sharing a link or two?
Familiar with MrB but no idea about anything blowing up. Quick search doesn't make the referenc obvious.
That’s the original video that kicked it off. It’s pretty mundane as far as internet controversy goes. The lotteries were run poorly and as such were probably illegal, not exactly uncommon on YouTube. The show largely targets kids, so there’s some accusations of it effectively being gambling for kids.
After that people dug into his whole life and found out one of his employees was convicted of statutory rape and he was aware, which is a bad look for a channel mostly watched by children.
It’s like a 3/10 on internet controversy scales. Doing something common and dumb but not obviously malicious, and then just stuff that looks bad without concrete evidence anything bad actually happened.
It’s not terribly interesting, I only saw because of a comedian making fun of it.
I never get any relevant ads or videos for me :(.
It is absolutely possible to serve recommendations in a content agnostic way based on other user behavior. That is how the Youtube's recommendations worked at least initially. It should be obvious that they didn't have this ability to analyze content when Youtube started almost 20 years ago.
And even in a content aware system, there is a difference between matching up keywords and understanding what those words mean. That is why an automatic transcript feels more neutral. That just transcribes what is being said. There is no implication of any analysis on the meaning of what is being said.
>If anything a dumb rule like "highlight the subscribe button if the transcript mentions it" is much more benign then a recommendation engine that can somehow always serve you a relevant video.
I'm not saying this feature is scarier than the recommendation engine. I'm saying Google revealing how much they know about the content of the video makes everything they do, including their recommendation engine, a little scarier.
CSS animations werent there I guess, but given a manual transcription, either by the submitter, or by a person watching, youtube definitely had the capability to do exactly this, maybe with the button being some flash animation.
I mean the "given a manual transcription" condition means the very much didn't have "the capability to do exactly this". The animation itself isn't the interesting part here.
Yes but the quality and accuracy of that system would be inferior, with a long delay until it gathers enough collaborative filtering information to be useful.
Indeed, and I assume that's why the home page is full of sports and music when I'm not logged in. (Or was recently, it now shows a blank page).
I'm a very strange person: I don't have any interest in sports, and will listen to perhaps 2 or 3 bits of music every week — if I've heard a piece before, I generally find it predictable and by extension dull, unless some prior association of positive feelings can override that.
1) similarity to other users (people who watched this also watched)
2) similarity in words in title or description (or the generated transcript)
I honestly suspect this feature is just implemented as a side effect as the auto-captioning... and now I'm a bit curious if carefully crafted manual captions can trigger it to go off continuously.
How did you think they did automated closed captioning..?
Think of it like reading a foreign language that uses the same alphabet. You can give me some French text and I can read it well enough for a French speaker to understand it. That doesn't mean that I myself understand what I read. Transcripts are the former and this feature is the latter.
I'd expect the button to light up at random times in a video about "smash the like button" causing the button to light up, rather than in the end call to action for the video where the video author actually wants the watcher to click the button.
Similarly, I doubt it can properly understand when the video presenter implies the term "smash the like button" while instead say, leaving an empty space where they would usually say such a thing.
If I pointed you at some french text and asked you to point at the term "le Chateau" you'd be able to do that without understanding french.
> It's just a substring match.
Let's just say this is true. That is a super simply process, but what would it look like?
Step 1: Transcribe the audio into text
Step 2: Run substring match on text
The transcribing/close captioning feature only does step 1. This shows that a step 2 is possible. I think you would have to be naive to think the capability to do this type of analysis on the transcribed text was designed for only this feature and would never be used for anything else. This feature is announcing that Youtube isn't merely creating transcripts of the audio in videos, it is running some unknown amount of analysis on that data.
As I said in my original comment, "it isn't that I didn't know Google had this ability", but this is literally a glowing sign pointing to this fact. I think the danger of Google reminding people of this has potential to outweigh the benefit of the "that's cute" reaction that this is designed to elicit.
Why would you have to be naive to believe this? The subtitles, with timings, are available on the client-side already. You seem to allude that doing this work would require some sort of deep analysis work. I think it's really more like 5 lines of JS, and 4 of them are producing the fun animated gradient :P
This feature is like walking into your kitchen one day to find a dead cockroach. I’m saying that is an indication of a cockroach problem while you’re effectively responding with “it’s fine, the one cockroach is dead and there is no reason to believe there are any others.”
> This feature suggests Google understands what was said.
Running .includes() on the client does not imply that Google has any "understanding" of what was said. It only implies that they ran .includes() on the client. includes() does not "understand" anything.
The thing I really don't understand is this: the fact that Google has closed captions at all implies they do an enormous more "analysis" than this minor feature could possibly require. If you understand how Google does CCs and what that means, this shouldn't have bothered you at all.
In your analogy, it's like you see a mound of a hundred thousand cockroaches, but you're are worried about a dust speck in another room.
Do you want to have a good faith conversation about this? Because going back to debating the meaning of "understanding" after I already said this was misleading is not a good way to have a conversation.
>The thing I really don't understand is this: the fact that Google has closed captions at all implies they do an enormous more "analysis" than this minor feature could possibly require. If you understand how Google does CCs and what that means, this shouldn't have bothered you at all.
Can we set up a baseline that there is a difference between content agnostic analysis and content aware analysis? Transcripts are content agnostic in that they can be produced without any comprehension of the words said. This feature is content aware in that it is looking for specific meaning in the words said. Do you not see any difference between these two?
Regardless, there is in my opinion a clear distinction in sophistication between a filter and something that triggers a timed action. And that was really what my original comment was about, this feature's elevated sophistication is a conscious reminder of Google's capabilities. Normally that is out of sight and out of mind which is probably better for Google.
Call it "understanding", call it "content aware analysis". I guarantee that their closed captioning service has much more of that quality than this new feature does.
> Can we set up a baseline that there is a difference between content agnostic analysis and content aware analysis? Transcripts are content agnostic in that they can be produced without any comprehension of the words said. This feature is content aware in that it is looking for specific meaning in the words said. Do you not see any difference between these two?
Again, I don't see it. CCs are not content agnostic: they have to have semantic understanding of the words said in order to produce accurate results. How do you think CCs differentiate between the words "to", "too" and "two" without looking at the surrounding words and having some idea of contextual usage? How do you think CCs can tell between "there" and "they're" without understanding if the speaker is referring to a person or a location? This is only the tip of the iceberg as to how CCs actually work, and more "content aware analysis" will always lead to more accurate CCs.
Still can't get away from that "understanding" debate. You're also now equating an understanding of context with an understanding of meaning. An understanding of meaning isn't needed to differentiate between "to", "two", and "too" because they're all used differently in sentences. When the system encounters those, I don't think it goes to the definitions and tries to find which word makes the most meaningful sentence. Most times the specific homophone can be inferred based on things like part of speech and the part of speech can often be inferred from a sentence without knowing any meaning.
For example, would the system be able to properly handle homophones that are grammatically similar? Could it consistently transcribe sentences like "I have Celiac disease and enjoy the taste of rose water, so I prefer flower to flour in my deserts." That is an easy sentence to understand for anyone who knows the meaning of those words, but there are no grammatical or structural indications on which flower/flour to use.
But either way, that is getting way too deep in the weeds compared to where my point started. This feature calls attention to an analysis of meaning because the user sees the software reacting to the meaning of the content of the video. A transcript does not call attention to an analysis of meaning because the behavior of the software does not change based on the content of the video.
Your first comment - the one that started all this - was, as far as I can understand, arguing that this feature indicated that Google had the capabilities to do more advanced - understanding? processing? meaning analysis? - than it had done in the past. If I keep coming back to that, well, it's because it appears to be your main point. If it's not, please correct me.
> Most times the specific homophone can be inferred based on things like part of speech and the part of speech can often be inferred from a sentence without knowing any meaning.
This is not true. I don't think I have enough more responses on HN to fully explain why homophones can not be inferred without understanding meaning, but I encourage you to go and read about how transcription works!
> For example, would the system be able to properly handle homophones that are grammatically similar?
I mean, this is easy enough for you to check. Here's some videos about flour / flower - notice how the CCs correctly determine if the word is flour or flower with almost 100% accuracy.
https://www.youtube.com/watch?v=y8vLjPctrcU https://www.youtube.com/watch?v=xdaRvErv2Kc
> This feature calls attention to an analysis of meaning because the user sees the software reacting to the meaning of the content of the video.
Are you saying you specifically think that YT is analyzing meaning from this feature, or just some generic user? I think you are smart enough to know that it's not true, but perhaps my mom might not understand that CCs require infinitely more processing power and this feature is just a drop in the bucket. (If you really still don't think it's true, definitely go read more about how CCs are made!)
Here is what I said. "It highlights how much Google analyses the content of its videos... It isn't that I didn't know Google had this ability...". My point was not that I learned about Google's capability from this feature or that this capability was new, it is that this calls attention to Google looking for meaning in the content of the video. A transcript does not call attention to Google looking for meaning regardless of how the transcripts are prepared.
>I mean, this is easy enough for you to check. Here's some videos about flour / flower - notice how the CCs correctly determine if the word is flour or flower with almost 100% accuracy.
Both those videos include the correct homophone in the title and description of the video. Choosing the correct one is not an indication of the system using the meaning of those words, it is pattern recognition. Every use of "flower" means the next usage is less likely to be "flour". The specificity of the example I used was important because it used both "flower" and "flour" in a way that can only be distinguished by the meaning of the words.
>Are you saying you specifically think that YT is analyzing meaning from this feature, or just some generic user? I think you are smart enough to know that it's not true, but perhaps my mom might not understand that CCs require infinitely more processing power and this feature is just a drop in the bucket. (If you really still don't think it's true, definitely go read more about how CCs are made!)
This feature is a glowing sign that Youtube as a company analyses the content of the videos for the meaning of what is said in those videos. You are too deep into the technical details trying to assign credit for what aspect of Youtube does the "understanding" or which "require[s] infinitely more processing power".
Think of this feature like receiving mail and you see one of the letters has already been opened. That could make you feel like your privacy was invaded in a way you wouldn't feel after receiving a postcard. And now we have spent several comments debating whether a torn envelope indicates whether anyone read the letter and whether a postcard is private.
Try 1 to 2.
I mean, it has autogenerated captions, translated (which i use to learn languages, fun)
The biggest issues tend to be when "Big Brother" style companies perform that kind of analysis on your content like the Microsoft AI screenshotting backlash.
I think I've missed this - what happen?
So that sweet engagement KPI is positively impacted and Youtube's CRO/experiment/Ab testing team gets their cookies.
There are probably various small tweaks like this in the UI.
Google analyses the content of the uploaders' videos yet "YouTube search", from the end user's perspective, is seemingly searching strings in titles, maybe tags. These titles generally suck for this purpose, even containing unicode garbage. Throw in all the "recommended" crap in the "results", videos that do not contain the strings being searched, and it barely feels like search at all. The end user is sent "results" she never asked for.
There's a whole bunch of lore about the types of words and phrases (such as slurs and profanity) that can affect algorithm rankings, monetization, or child-restriction.
Doing a regex search for "smash that like button" is trivial by comparison.
From a technical perspective I like the attention mechanism and that it's automated.
(I use the VTT on my Firefox, I use IDM to download the auto-generated English language scripts for lengthy videos).
I remember listening to Chomsky saying that he prefers to read the transcripts than listening/watching speeches.
It is one of those things I don't care (at all) to know how it is technically set up :)
The best calls to action inside a video take advantage of the Zeigarnik effect : once they're invested in the content, people really want to find out the
I turned the TV off and walked away...
> Davie504 in pieces
I think there’s a similar effect for the subscribe button. But I don’t remember which phrase triggers it
It might even just be conventional speech-to-text, which can struggle with accents and poor acoustics, etc.
No need for fancy AI, on top of whatever technology powers the transcription feature.
I’m wondering what an unbiased survey would even look like.
Maybe I have just become more attentive to these things or they have become more prevalent, but I notice this blatant optimization for the Algorithm(tm) more and more. I guess this is the meta you need to play in order to stand out/make it in the torrent of content that is YT or social media in general. I also see this as an example of Goodhart's law as creators try to optimize for reach and thus likes etc. while the original intention of the YT Algorithm(tm) was (or should have been) to serve you content you will enjoy.
The best example of this is clickbait thumbnails. Nearly every creator I am subscribed to has expressed hatred for them, and yet they mostly still make them. Why?
Because, when youtube shows your thumbnail to someone, if they do anything other than click on that thumbnail, your channel/video is downranked in the algorithm (because it has "low conversion"). Every creator is in a zero sum, fully antagonistic war with every other creator to get you to click on their video instead of someone else's. If a channel has low enough "conversion" for long enough, youtube will stop showing that video to potential viewers, including your actual subscribers. For a channel to survive and make enough money to pay rent and things, you have to play the dumb algorithm games. Patreon is literally the only alternative.
Youtube prefers it this way because then they get to crowdsource more addictive and "engagement driving" content through millions of creators, rather than just a department in google (which honestly they probably still have). They didn't have to tell anyone to make "more stupid clickbait thumbnails", they did it themselves because it was encouraged by the incentives given to them. A Youtube where everyone plays the algorithms and thumbnail games literally earns more money and views than a youtube where every creator is only driven by their own desires to make good art. Google has been doing this since they bought Youtube. Early on, only videos longer than ten minutes could get ads, so every video ballooned to ten minutes and one second in length.
I'm sure many here would love to make use of some uBlock Origin rulesets to disable them?
Before Animation: <yt-smartimation class="smartimation smartimation--experiment-enabled smartimation--enable-masking">
While Animation: <yt-smartimation class="smartimation smartimation--experiment-enabled smartimation--active-border smartimation--active-background smartimation--enable-masking">
Yet...they can't hire more staff, implement databases to track malicious/unreliable DMCA submitters, and falsely-targeted creators for whom complaints are likely BS, and implement AI-powered tools to assist with all of this?
Example: Over here there is a huge problem with rented e-scooters being left wherever.
Every scooter company asks you to take a photo of the parking. NOBODY ever checks those.
I did a quick test with GPT-4 and it could tell me if a scooter was parked in a considerate way about 90% of the time (10 photos, some with multiple scooters).
Could they implement this in an afternoon? Yes. Will they? No. Because it'd be used to penalise their own customers.
Only a semi-scooter company but Lime, at least in London, do sometimes check the photos because I once got dinged for "incorrect parking" (wrongly, mind, it was parked identically to the Santander bikes a few metres away and thus incapable of inconveniencing anyone.)
The thing that I do see on every video that reminds me to hit the like button is seeing the counter move. I'm way more likely to see that and think "yeah I do like this video, I should hit this button" when I see the counter change than I am to have a similar reaction when explicitly told by the person in the video to do it.
I'm sure thay's not ideal (e.g. not for Google trying to understand you better) for the addicts, but for the other ones, this can be the perfect solution.
I need to balance all the time spent on this site reading about using Kubernetes at scale, AI changing the world and how to bring more value to my employer with their videos.
I don't blame creators for doing this: creators have to say this to compete (this is an example of Moloch[1]). The point isn't to hurt creators who do this, it's to remove the incentive so that creators no longer have to do this. So I think the proper way to do this is to announce that videos that ask for likes or subscribes will be penalized, and only penalize videos that come out after the announcement.
But, YouTube serves advertisers, not viewers, so they'll never do anything that benefits viewers but hurts advertisers.
[1] https://slatestarcodex.com/2014/07/30/meditations-on-moloch/
[1] https://www.nrk.no/ostfold/xl/tiktok-doesn_t-show-the-war-in...
So, at first I thought that Google is parsing the text of the video being played, but is not listening to your mic.
However, several comments both here and on YT imply that Google is ALSO listening to what you say on your mic. Which, frankly, is much more scary - though I cannot currently test.
Other than that, why wouldn’t they know what’s being said in a video they host?