How big is YouTube?
ethanzuckerman.com
ethanzuckerman.com
What makes it the "best method"? Would it be better to use a seine, or a trap, or hook-and-line? How would we know if there are subpopulations that have different likelihood of capture by different methods?
To say it's "reasonable to argue that the effect is insignificant" is purely assertion. Why is it unreasonable to argue that a fish could learn from the first experience and be less likely to be captured a second time?
If what you mean is that it's better than a completely blind guess, then I'd agree. But it's not clearly the best method nor is it clearly unbiased.
And, also, the fish gets tagged for its own good.
They'll even call you tinfoil fish.
In the "100 fish" example, the formula for approximating the total number of fish is:
total ~= caught / tagged
(where caught=100 in the example)
In their YouTube sampling method, the formula for approximating the total number of videos is: total ~= (valid / tried) * 2^64
Notice that this is flipped: in the fish example the main measurement is "tagged" (the number of fish that were tagged the second time you caught them), which is in the denominator. But when counting YouTube videos, the main measurement is "valid" (the number of urls that resolved to videos), which is in the numerator.Edit: Oh because 64^10 * 16 = (2^6)^10 * (2^4)
You could bucket bugs into categories by severity or type and that might improve the estimate, as well.
I guess it underestimates the number hard to find bugs though since it assumes same likelyhood to be found.
https://en.wikipedia.org/wiki/Unseen_species_problem http://www.stat.yale.edu/~yw562/reprints/species-si.pdf
Won't this mess up stats though? It's like a lake monster randomly swapping an untagged fish with tagged fish as you catch them.
We need to figure out how to target the malicious individuals and groups instead of getting creeped out by them to the point of destroying most of the so praised democratizing of computing. Between this and locking down the local desktop and mobile software and hardware, we've never got to having the promised "bicycle for the mind".
If people had perfect self-control, they wouldn't do it. IMO it's somewhat irresponsible for the algorithm makers to profit from that - it's basically selling an unregulated, heavily optimized drug. They downrank scammy content for instance, which limits its reach - why not also downrank trolling? (obviously bc the former directly impacts profits, but not the latter, but still)
You can legitimately claim that people respond in a very striking and predictable way to being set on fire, and even find ways to exploit this behavior for your benefit somehow, and it still doesn't make setting people on fire a net benefit or a service to them in any way.
Just because you can condition an intelligent organism in a certain way doesn't make that become a desirable outcome. Maybe you're identifying a doomsday switch, an exploit in the code that resists patching and bricks the machine. If you successfully do that, it's very much on you whether you make the logical leap to 'therefore we must apply this as hard as possible!'
With the CA scandal, now all the big companies would lock down their app data and sell ads strictly through their limited API only, so the ads buyer would have much less control before.
It's basically saying: u cant behave with the open data. Then we will do business only
Companies are locking down access to public posts. This has nothing to do with CA, just with companies moving away from the open web towards vertical integration.
Companies requiring users to login to view public posts (Twitter, Instagram, Facebook, Reddit) has nothing to do with protecting user data. It's just that tech companies now want to be in control of who can view their public posts.
Isn't this ironic, given how google bots scour the web relentlessly and hammer sites almost to death?
I'm biased against Google, but I'm honest about it. I don't ask "what could I possibly be biased about?"
I have been hosting sites and online services for a long time now and never had this problem, or heard of this issue ever before.
If your site can't even handle a crawler, you need to seriously question your hosting provider, or your architecture.
If your site is very popular and the content changes frequently, you can find yourself getting crawled a higher frequency than you might want, particularly since Google can crawl your site at a high rate of concurrency, hitting many pages at once, which might not be great for your backend services if you're not used to that level of simultaneous traffic.
"Hammered to death" is probably hyperbole but I have worked with several clients who had to use Google's Search Console tooling[0] to rate-limit how often Googlebot crawled their site because it was indeed too much.
0: https://developers.google.com/search/docs/crawling-indexing/...
also for less friendly crawlers a rate limiter is needed anyway :(
(of course the existence of such tools doesn't give carte blanche to any crawler to overload sites ... but let's say they implement some sensing, based on response times, that means a significant load is probably needed to increase response times, which definitely can raise some eyebrows, and with autoscaling can cost a lot of money to site operators)
My point was that people provision capacity ideally based on observed or expected traffic, and that crawlers can, and do, show up and exceed that capacity sometimes, having a negative effect on your customers' experience.
But you are correct that it's absolutely manageable. And telling crawlers to slow the F down is one of the tools you can use to manage it. :-)
That's pre-cloud ubiquity so scaling up meant buying servers, installing them on a data center, and paying rent for the racks. It was a fucking nightmare to deal with.
I expect the optimal solution is to increase the address space to prevent random samples from collecting enough data to arrive at a statistically significant conclusion. There are probably other good solutions which attempt to vary the distribution in different ways, but a truly random sample should limit that angle.
Take phone numbers as an analogy. Without foreknowledge or without careful analysis of the distribution of phone numbers, you might assume all numbers are valid, but in fact 555-xxxx is always invalid within each area code [0]. For each set of reserved numbers our address space is that much smaller, which can skew the results of the statistics we gather from it if we don't exclude them from our original calculations.
It may be that YouTube reserves off certain address spaces (eg maybe it can't start with a 0, or maybe two visually similar values cannot be next to each other (eg I and 1), etc), which may make this sampling method (slightly less) accurate than it might otherwise appear.
Suppose I have addresses /v=0x00 to 0xff, but I only use f0 to ff; if you assume the videos are distributed randomly then your estimates will always be skewed, no?
So I take the addressable space and apply an arbitrary filter before assigning addresses.
Equally random samples will be off by the same amount, but you don't know the sparsity that I've applied with my filter?
There is no way for clustering to alter the probability of a hit or a miss. There is nowhere to "hide". The probability of a hit remains the proportion of the space which is filled.
(it is named this way because it was an archival effort to collect the information before the dislike feature was removed)
It can be used to find things like the most controversial videos; top videos with a description in a certain language, etc.
from the article:
> It’s possible that YouTube will object to the existence of this resource or the methods we used to create it. Counterpoint: I believe that high level data like this should be published regularly for all large user-generated media platforms. These platforms are some of the most important parts of our digital public sphere, and we need far more information about what’s on them, who creates this content and who it reaches.
The gov't ought to make it regulation to force platforms to expose stats like these, so that it can be collected by the statistics bureaus.
So has been the big banks, large corporations, land but they all feed off each other and the government. What we want as a community is usually quite different to what they decide to do.
Nobody is stopping users from selecting one of the many YouTube competitors out there (eg - Twitch, Facebook, Vimeo) to host their content. We could also argue that savvy marketers/influencers use multiple hosting platforms.
YouTube's data is critical for YouTube and Google, which is basically an elaborate marketing company.
Governments should only enforce oversight on matters such as user rights and privacy, anticompetitive practices, content subject matter, etc.
With very limited exception, Vimeo imposes a 2TB/month bandwidth limit [0] on all accounts. If you exceed that limit and don’t agree to pay for your excess usage, Vimeo will shut you down.
[0] https://help.vimeo.com/hc/en-us/articles/12426275404305-Band...
I got 400KB/s from some FHD 24-30 fps videos I downloaded, but this is very approximate. YouTube will encode sections containing less perceptible information with less bitrate, and of course, videos come in all kinds of different resolutions and frame rates, with the distribution changing over the history of the site. If we assume every video is 4K with a bitrate of 1.5MB/s, that's 10 exabytes.
This estimate is low for the amount of storage YouTube needs, since it would store popular videos in multiple datacenters, in both VP9 and AV1. It's possible YouTube compresses unpopular videos or transcodes them on-demand from some other format, which would make this estimate high, but I doubt it.
Pretty wild stuff.
400KB/s, or 3.2Mbps as we would commonly use in video encoding, is quite low for original quality upload in FHD or commonly known as 1080p. The 4K video number is just about right for average original upload.
You then have to take into account YouTube at least compress those into 2 video codec, H.264 and VP9. Each codec to have all the resolution from 320P to 1080P or higher depending on the original upload quality. With many popular additional and 4K video also encoded in AV1 as well. Some even comes in HEVC for 360 surround video. Yes you read that right. H.265 HEVC on YouTube.
And all of that doesn't even include replication or redundancy.
I would not be surprised if the total easily exceed 100EB. Which is 100 (2020 ) Dropbox in size.
No, it definitely is not.
I pine for the day when "hella-" extends the SI prefixes. Sadly, they added "ronna-" and "quetta-" in 2022. Seems like I'll have to wait quite some time.
(from https://iopscience.iop.org/article/10.1088/1681-7575/ac6afd, via wikipedia)
Seems like we missed opportunity to have an official metric fuckton.
Hyundai Kona on the other hand was way more serious and they changed it to another Island in the Portuguese market. Kona's (actual spelling "cona") closest translation would be "cunt", in the US sense in terms of seriousness, not the Australian more light one.
Source is I'm portuguese
Rats. So close!
IIUC, ~1ZB is basically the entire hard drive market for the last 5 years, and drives don't last that long...
I suspect YouTube isn't >10% of all Google.
On the other hand: there might be a lot of videos with ridiculously low view counts.
On the third hand: remember that YouTube had to come up with their own transcoding chips. As they say, it's complicated.
Source: a decade ago, I knew the answer to your question and helped the people in charge of the storage bring costs down. (I found out just the other day that one of them, R.L., died this February... RIP)
The lots of videos with low view counts are accounted for by the article. It sounds like the only ones not included are private videos, which are probably not that numerous.
Granted, none of them are uncompressed 4K terabyte sized files. I haven't got my originals to do a bit-for-bit comparison. But judging by the filesizes and metadata, they are all the originals.
source interviewed by Google a few times
> Jason found a couple of cheats that makes the method roughly 32,000 times as efficient, meaning our “phone call” connects lots more often
Proving that using cheats and auto complete does not break sample independence and keeps sampling as random as possible would be needed here for stats beginners such as me!
Drunk dialing but having a human operator that each time tries to help you connect with someone, even if you mistyped the number... Doesn't look random to me.
However I did not read the 85 pages paper... Maybe it's addressed there.
> By constructing a search query that joins together 32 randomly generated identifiers using the OR operator, the efficiency of each search increases by a factor of 32. To further increase search efficiency, randomly generated identifiers can take advantage of case insensitivity in YouTube’s search engine. A search for either "DQW4W9WGXCQ” or “dqw4w9wgxcq” will return an extant video with the ID “dQw4w9WgXcQ”. In effect, YouTube will search for every upper- and lowercase permutation of the search query, returning all matches. Each alphabetical character in positions 1 to 10 increases search efficiency by a factor of 2. Video identifiers with only alphabetical characters in positions 1 to 10 (valid characters for position 11 do not benefit from case-insensitivity) will maximize search efficiency, increasing search efficiency by a factor of 1024. By constructing search queries with 32 randomly generated alphabetical identifiers, each search can effectively search 32,768 valid video identifiers.
They also mention some caveats to this method, namely, that it only includes publicly listed videos:
> As our method uses YouTube search, our random set only includes public videos. While an alternative brute force method, involving entering video IDs directly without the case sensitivity shortcut that requires the search engine, would include unlisted videos, too, it still would not include private videos. If our method did include unlisted videos, we would have omitted them for ethical reasons anyway to respect users’ privacy through obscurity (Selinger & Hartzog, 2018). In addition to this limitation, there are considerations inherent in our use of the case insensitivity shortcut, which trusts the YouTube search engine to provide all matching results, and which oversamples IDs with letters, rather than numbers or symbols, in their first ten characters. We do not believe these factors meaningfully affect the quality of our data, and as noted above a more direct “brute force” method - even for the purpose of generating a purely random sample to compare to our sample - would not be computationally realistic.
In short I do believe that the sample is valuable, but it is not a true random sample in the spirit that the post is written, there is a heuristic to have "more hits"
That's very clever. Presumably the video ID in the URL is case-sensitive, but then YouTube went out of their way to index a video's ID for text search, which made this possible.
If you don't know how the checksum is created you can still try all values of it for one sample of the actual ID space.
So you issue an API to create a playlist with video IDs x, x+1, x+2, ..., and then when you retrieve the list, only x+2 is in it since it is the assigned ID.
> it was discovered by Jia Zhou et. al. in 2011, and it’s far more efficient than our naïve method. (You generate a five character string where one character is a dash – YouTube will autocomplete those URLs and spit out a matching video if one exists.)
1. Assume a range of values
2. Assume a fair probability function for sampling over the range of values
The estimated size is the %-of-hits * the total range of values.
1. So let's say that possible range of values is true (10 characters of specific range + 1). That would represent one big circle of possible area where videos might be.
2. Distribution of identifiers (valid videos) is everything. If Youtube did some contraints (or skewing) to IDs, that we don't know about, then actual existing video IDs might be a small(er) circle within that bigger circle of possibilities and not equally dispersed throughout, or there mught be clumping or whatever... So you'd need to sample the space by throwing darts in a way to get a silhouette of their skew or to see if it's random-ish, by I don't know let's say Poisson distribution.
Only then one could estimate the size. So is this what they're doing?
Also.. anyone bothered to you know, ask Youtube?
This is the risk of explaining the method
Video IDs are immutable
A video can only be represented by a single unique video ID
etc.
Presumably, no one would do that except researchers trying to count videos (or randomly find hidden ones?).
You can break the assumption of unicity (if an unassigned ID is later assigned) if you do that internally, although not sure that’d be common but it’s not an assumption that has to be strict for non-attributed ones, and you never use the fake ID.
if you got a video from a randomly generated ID, you can immediately query for it again, and see if the video was the same as before.
If it's not the same, you disgard the result and assume the ID generated was actually non-existent.
If it's the same, then you know it's a real ID.
As long as youtube vidoe URLs are immutable, this method will stop the blocks you described.
a lot of work, for very little gain.
I'm going to roll a d20, which gives me a uniformly distributed integer from 1 through 20. I'll look into the bin with that number and if there is something there take it.
You don't want me to take any of your things. How can you distribute your 3 items among those 20 bins to minimize the chances that I will take one of your items?
If I were not using a uniformly distributed integer from 1 through 20 and you knew that and knew something about the distribution you could pick bins that my loaded d20 is less likely to choose.
But since my d20 is not loaded, each bin has a 1/20 chance of being the one I try to steal from. Your placement of an item has no affect on the probability that I will get it.
It works the same the other way around. If you place the items using a uniform distribution, then it doesn't matter if I use a loaded d20, or even just always pick bin 1. I'll have the same chance of getting an item no matter how I generate my pick.
In general when you have two parties each picking a number from a given space, if one of the parties is picking uniformly then nothing the other party can do affects the probability of both picking the same number.
Now imagine a number of these bins _can hold zero things_ (not 3). Eg in a world where all bins are the same size, you can always steal 3 things from any of the bins, whereas in a world where the bin sizes vary. You'd hit a few bins which are guaranteed empty. Doesn't this directly affect the probabilities?
This is the best thing I've ever seen.
Why would they not be random? Nobody has ever found a pattern that I'm aware of, and there are pretty solid claims of past PRNG use. And a leak of the PRNG seed was likely why they mass-privated all unlisted videos a couple years ago.
If you believe that there is some structure in YouTube video IDs, that would have no effect on this experiment. It would just reduce the fraction of the total address space that YouTube can use. This is a well-known property of "impure" names, and it means there is a good chance that the IDs have no structure. In other words, the video IDs would be "pure" names.
They got 10,000 samples of hits, and a huge number of samples of misses. Their result should be very accurate. (32,000 was a different number)
I'm not sure how much the "cheating" would affect the precision of the result. But assuming it has no effect, it's easy to estimate this precision:
They found X = 24964 videos in a search space of size S = 2^64. For the number of existing videos they report the estimate N = 13,325,821,970. From this we can find their estimate for the probability that a particular ID links to a video: p = N / S ≈ 7.22e-10. So the equivalent number of IDs that they have checked (the number of checks without cheating that would give the same information) is n = X / p ≈ 3.46e13.
Since X is a Binomial, its variance is Var(X)=n⋅P(1-P) (where P is the real proportion corresponding to the estimate p above). And N = X⋅S/n so its variance is Var(X)⋅S^2/n^2. The standard deviation of N is thus σ = S⋅sqrt(P⋅(1-P)/n). Now we don't know P but we can use our estimate p instead to find an estimate of σ!
We find that the standard deviation of their estimator for the number of YouTube videos is approximately S⋅sqrt(p⋅(1-p)/n) ≈ 8.43e7. That's just 0.633% of N so their estimate is quite precise.
It would be even cooler if it had a deeper view of categories. Right now the biggest category by far is People & Blogs but it's possible to get much more information if it was broken down into sub-categories.
Last night they issued a notification to my phone requiring me to update the YouTube app.
The problem - it is the last version that runs on my phone.
At least web still works, for now.
Several is 3 to 5? So 4 months.
10 000 videos over 4 months. So 2500 per month. If we assume 20 work days.. that's 125 per day. Over 8 hours.. 15,6 per hour.
Those scripts seem to be quite slow? Was this done by hand? A video every 3 minutes.
Turns out there are 2^64 possible YouTube addresses, an enormous number: 18.4 quintillion. There are lots of YouTube videos, but not that many.
Let’s guess for a moment that there are 1 billion YouTube videos – if you picked URLs at random, you’d only get a valid address roughly once every 18.4 billion tries.
I mean, sure, they did reduce that a little: Jason found a couple of cheats that makes the method roughly 32,000 times as efficient, meaning our “phone call” connects lots more often.
With that in mind, how many attempts did they make to get a hit every three minutes?It's surprising that they weren't throttled back for making excessive requests.
> generate a random 5 character long string where one of them is a dash, and Youtube will autocomplete the string if a matching video exists
A random youtube video service would be interesting! Like Wikipedia random, but less educational XD
in a few days your home feed and recommendations will consist of random low-ranked junk. Then, start blocking the highest views channels and repeat the 1.2.3. steps to get even more obscure content. No need to scrape millions of IDs or flood the servers with random requests.
Can anyone post what this site is? So I can tell whether this #1 story is worth it trying a different ISP or user agent string or whatnot
In the meantime, this archive page should work for you: https://web.archive.org/web/20231223131348/https://ethanzuck...
Well, I click the link and get the access denied page. Do you need an IP address (93.135.167.200), user agent string¹, timestamp, or what would help?
Randomly noticed my accept-language header is super convoluted, mixing English-in-Germany, English, German in Germany, Dutch in the Netherlands, Dutch, and finally USA English, with decreasing priority from undefined to 0.9 to 0.4 in decrements of 0.1. I guess it takes this from my OS settings? Though I haven't configured it explicitly like that, particularly the en-US I'd not use because of the inconsistent date format and unfamiliar units system. Maybe the server finds it weird that six different locales are supported?
Thanks for responding and relaying!
¹ Mozilla/5.0 (Linux; Android 10; Pixel Build/QP1A.190711.019; wv) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/76.0.3809.132 Mobile Safari/537.36
LOL
Results are entirely off-topic, but related to other interests I have; tech, planes, trains, and automobiles, etc.
There are some weird corners of YouTube.
> NOT big enough; that it can’t be wiped away from the surface earth; with a relatively cyber attack from the Alliance. In fact; Google/YouTube will just be one; of 63 communist controlled platforms; that will be completely erased from existence once the order is given.
I must know these things now.
> NOT big enough; that it can’t be wiped away from the surface earth; with a relatively cyber attack from the Alliance. In fact; Google/YouTube will just be one; of 63 communist controlled platforms; that will be completely erased from existence once the order is given.
I have to wonder what this "Alliance" is supposed to be.
There are many variations on this theme, and though not widely held as a belief, elements seep into the wider discourse and thought.
I'm not sure what the aliens' motivation would be in this scheme, and often they're not part of the story.
EDIT: here's a random page that likely will bring more questions than answers but it references the elements described above ... https://divinecosmos.com/davids-blog/22143-moment-of-truth-q... ... there are variations of this conspiracy theory tailored to many different tendencies: New Agers, Evangelical Christians, White Supremacists, Conservative Republicans, etc. For a further look, find interesting interviews with a certain Jan Halper Hayes and her supposed work with USA Space Force. Some of these promoters a true believers, others are grifters. Also ... grand corporatists and communists are usually conflated in these narratives.
Something vaguely related to next year's A24 release "Civil War", but much more QAnon-adjacent and informed. Even better if Q would denounce it all as an attempt to "innoculate people from the truth".
In experiments by yours truly, it seems that quality goes down as views go up.
However, that's also an opportunity that YouTube hasn't realized yet, they could become a primary platform for content if they had a better algorithm and a better search, it would help them to monetize better those subscriptions.
Youtube has one of the lowest per viewer revenue and the poor discovery isn't an accident in that.
Are you sure you live so far from that world?
My experience of YT is totally different, FWIW.
It's YouTube which thinks there's only 10 videos worth watching, not me.
The average length of a youtube video is around 13 minutes. My napkin GPT math values video per minute at $35 per minute.
My napkin GPT math values each video closer to $455.
This leaves me with one unanswered question… Where’s the beef?
I've read that the number of views / 100 in dollars is a good estimate of how much the channel receives.