Gangnam Style breaks YouTube viewer count
plus.google.com
plus.google.com
I made a bug once, and I need to tell you about it. So, in 2001, I wrote a
reference library for JSON, in Java, and in it, I had this line
private int index
that created a variable called "index" which counted the number of characters in
the JSON text that we were parsing, and it was used to produce an error message.
Last year, I got a bug report from somebody. It turns out that they had a JSON
text which was several gigabytes in size, and they had a syntax error past two
gigabytes, and my JSON library did not properly report where the error was — it
was off by two gigabytes, which, that's kind of a big error, isn't it? And the
reason was, I used an int.
Now, I can justify my choice in doing that. At the time that I did it, two
gigabytes was a really big disk drive, and my use of JSON still is very small
messages. My JSON messages are rarely bigger than a couple of K. And — a
couple gigs, yeah that's about a thousand times bigger than I need, I should be
all right. No, turns out it wasn't enough.
You might think well, one bug in 12 years you're doing pretty good. And I'm
saying no, that's not good enough. I want my programs to be perfect. I don't
want anything to go wrong. And in this case it went wrong simply because *Java
gave me a choice that I didn't need, and I made the wrong choice*.
[0] https://www.youtube.com/watch?v=bo36MrBfTk4&t=38mEDIT: is there a reference for formatting comments? I've never been able to find one.
*Java gave me a choice that I didn't need, and I made the wrong choice*If you're doing compression you should use some sort of raw bytes type.
At least I think that's the platonic ideal.
As long as you are concerned with adding a bit of eye candy and interactivity to a web page this may be true enough to get away with the JavaScript way of making every number a double precision floating point number but there a other domains where this will not fly. And even in the world of JavaScript asm.js is trying hard to overcome this limitation.
Sure, and I believe everybody, including Doug, knows this not so subtle distinction. He wasn't talking about programming HPC for NASA.
Most programs written in the real world (enterprise-y Java apps) do not need strong control on GC, choice of integer types, or many other things offered to them. Reducing choice will increase code/tool quality.
I think that we should make the uncommon choice reallllly hard to put into place. Make it a pain to configure the GC, give specific integer types really long names. Just stop people from premature optimization and leave these tools to people who know what they're doing.
Data manipulation beyond "pull out of database" or "submit user input to database" is a lot rarer in enterprise software like this than in scientific computing. I'm not saying it's bad to be aware of it, but software is more than numbers.
What was originally compared was different sized integral types.
What you are comparing is two numeric types that have a large semantic gulf (fractions vs integers).
So your point is disingenuous in the context of the former.
When it comes to integers you are right - signed and 32 bits is a viable choice in north of 90 % of all cases. And when I wrote you should be able to make good decision about the number type to use I was already thinking of all the number types, however I did not express this well. But then I really don't see a lot of difference between being able to choose between integer, floating point and decimal types on the one hand and various integer types on the other hand.
I do think that the difference between Integer/Fractional is important, but honestly if you're dealing with money you should be using some Money datatype that's smart about this instead of raw numbers.
For example, using Java's binarySearch on arrays of length over 2^30 was broken until 2006[0][1].
[0]: http://googleresearch.blogspot.com/2006/06/extra-extra-read-...
[1]: http://bugs.java.com/bugdatabase/view_bug.do?bug_id=5045582
Doesn't sound very humble, this "HO".
Do you know how many programmers that run circles around you (and me) have done mistakes similar to what Doug describes?
And how many web developers don't know about or are unable to decide between different number types? JavaScript type coercion is such a mess, how could you get away without thinking about types, even though number types are usually not an issue?
[1] http://people.mozilla.org/~jorendorff/es5.html#sec-11.9.3
[2] http://strilanc.com/visualization/2014/03/27/Better-JS-Equal...
`"Infinity" == Infinity`, but `"true" != true`.
I could go on for a while like this.
Language-war disclaimer: I love javascript and everything, it's a very expressive language; but a good wodge of the tooling around it nowadays is to help people avoid things like implicit type-coercion 'surprises'.
Other operators in JavaScript may behave more intuitively than ==, but I don't think you can really make a good case for JavaScript's type coercion being 'unsurprising'.
The second line at least uses type coercion, but still you are making the wrong assumptions. true could be coerced to many strings 't', '1', 'true', 'yes', 'on', but they chose to use '1' (You may not like it but I think it's a good choice). Infinity on the other hand has not many choices when coercing it to a string I can think of '∞' (which is difficult to type), 'Infinity', and maybe 'Inf.' so I think they made a good choice here.
I'm not saying type coercion in js has no problems, but you said that it's a mess and I just think you chose the wrong examples.
Anyhow, I don't agree with you: I think that getting a bug years down the road due to a too small numeric type is something that the programming language itself should avoid.... not because "developers don't ought to know it", but because mistakes happen
Anyhow, even with a dynamically typed programming language like Python or Javascript you can care about the size of your numbers.
Just import the array module (in Python) or use the Int32Array/Int8Array/etc. types (in Javascript)
What if Java gave an arguably more useful choice--whether to use a signed or unsigned integer?
Why is it even a question? Imagine a JSON of all the people on Earth.
And I certainly don't believe there are programmers making only 1 mistake for 12 years. I believe he's just making a joke, or using the example as a means to an end.
One the one hand "implicit int" are being phased out in favour of explicit int size. Your variables cannot be `int` anymore, you have to sit down, think and choose: u8? u16? u32? i32? u64? i64? This avoids all the pains of programs behaving differently or crashing when compiled on different architectures.
On the other hand, a new "native integer for sizes that do not matter, minimum 32 bits" is being brewed, for example for pointer offsets or collection sizes. The idea is that you will not be able to have a collection with more of 2^32 elements in a 32-bit architecture nor more than 2^64 in a 64-bit architecture.
After this discussion, my hope is to see the introduction of a fast-ish dynamic bigint (that starts native and grow up to 256 or 512 bits) that can be used in all the cases where you do not care about the exact size type, yet you want to be future-proof (this `private int index` fits this case, IMO).
[1] http://discuss.rust-lang.org/t/if-int-has-the-wrong-size/454... [2] https://github.com/rust-lang/rfcs/pull/464
It's only important for things like pointers that they match the size of the addressing space, and not even that is a very hard constraint, just a very convenient one.
However, using a 32 bit number to store values that will never be larger than 16 bits isn't that bad, it's just very slow.
However, an overflow to floating point isn't necessarily an improvement because, while a float will hold bigger numbers, it does so with limited precision and sometimes that lack of precision will cause bugs too. Probably more often, in fact.
In the example given it wouldn't be so bad, but you'd only get an approximate indication of where the error occurred rather than a specific line/character. So of course, whatever is reporting the error would now need to understand and handle the much more complex scenario of "fuzzy" location information instead of a simple unique index to a specific character. Depending on what it then needs to do with that information, the complexity could spiral from there.
If you want to just have things work no matter what, you have no choice but to use bignums. I was wondering about this recently, so did some benchmarks in Clojure. The performance was horrible, so frankly this is still not a viable alternative. Maybe in 10 years time, if every CPU has a bignum coprocessor by then.
Also, there are times, particularly in low-level graphics programming or cryptography, where you actually want integer modulo arithmetic, or to be able to do bitwise booleans predictably. In those cases, JavaScript-style loose typing can be a huge pain.
BTW, I've been a big advocate of JavaScript for about as long as Doug Crockford, so my point isn't that JavaScript-style type handling is bad: just that it's very far from a silver bullet.
========
What kind of formatting can you use in comments?
http://news.ycombinator.com/formatdoc
========
Was that what you were looking for?
IMHO Java makes some choices about safety. If you don't agree with those choices, use a different tool. It doesn't make Java wrong for having a different opinion. Likewise I wouldn't berate C for being too low level or Ruby for favouring readability over performance.
Put in perspective that is probably in excess the number of times the most favored "I Love Lucy" show has been seen. Or put another way, you've got a music video with the same eyeball impact as the highest rated television show ever.
That says to me that either advertising on Youtube is a bargain or advertising on TV is way over priced :-)
[1] http://tvbythenumbers.zap2it.com/2014/02/10/the-walking-dead...
(Of course, none of these streams show the same ads the original broadcast does—but if you're a clever ad agency, you're already doing product-placement instead of interstitials most of the time anyway.)
Can you name any examples of this?
The closest thing I can think of is Veronica Mars which was Kickstarted many years later and raised ~5 million dollars from 91,000 backers to make a single movie.
I think perhaps the "alternate consumption streams" viewers are not as lucrative as you think.
That everyone regretted backing.
The movie may not be good (my personal opinion), but clearly it met expectations.
Family Guy also had a similar fate, but not because of a botched launch, but because its audience existed, but did not consume television through mainstream sources. It was canceled after 2.5 seasons and then went on to become the best selling animated DVD series. Fox brought it back the next year.
I think you just backdoored into the most interesting ad campaign ever: 1) Find a show with a directory / writer / production team known for producing content that "stands the test of time" (e.g. likely to have a high total_views_over_time:broadcast_views ratio) 2) Include product placement for a non-existent product by a currently-existing company with strong brand recognition 3) Test response to non-existent product by initial viewers 4) Start viral campaign around non-existent product (this likely favors "Hunh?" shows a la Lost or Fringe) 5) Trigger view bump in show (win award, produce new episodes in partnership with Netflix, produce new movie, etc.) 6) Launch real-product multiple years after initial product placement
>there are plenty of other things to do on a computer while you wait through the ad.
Yea, for example you can buy the advertised product with a few clicks. If you are quick enough then you can finish the buying even before the video ad finishes (sure it's not the most realistic scenario but it's possible). Or with a quick search you can learn more about the product to check how honest is the ad. TV ads cannot compete with this efficiency. The only thing TV ads can do better is reaching bigger and the less tech interested audience.
For example, with I Love Lucy, the audience member likely sat and watched the entire commercial. With a YouTube video, the audience member can skip the ad or move on to other content.
TV = 22 minutes of content.
YouTube Video = 3 minutes of content.
Plus, the metrics that constitute views between the two media formats are completely different.
NEITHER. Advertising on YT is worthless and yes it used to be way overpriced on TV.
This song has been played many times by a relatively small section of the Dutch population :)
EDIT: Which is the solution they apparently implemented, converting signed to unsigned at some higher layer.
... and now I realize you may be correct, and that it's probably inevitable that a viewcount will not only exceed the total number of people alive, but will double or even quadruple it. Our total population is actually about 100 billion, but only ~7% of us are still alive.
The shadows of the dead will be forever enshrined as YouTube view counts. Our shadows.
"We always overestimate the change that will occur in the next two years and underestimate the change that will occur in the next ten." - Bill Gates
[One problem with unsigned integers in protocol buffers is that some supported languages like Python have no concept of signed/unsigned integers. Which is why for example Thrift does not support that distinction. This has nothing to do with C++ though.]
You can still implement range checks to enforce that the numbers are in the correct domain. It's not that much of a problem.
This is more of a problem in a statically typed language like Java, because it means there is no native data type to map protobuf's unsigned types to. In Python this doesn't matter as much because numbers will automatically promote themselves to bigints.
Edit: Now I realize that would mean Google couldn't have made this joke. But I am still not sure this was foreseen by Youtube devs from day one.
"You should not use the unsigned integer types such as uint32_t, unless there is a valid reason such as representing a bit pattern rather than a number, or you need defined overflow modulo 2^N. In particular, do not use unsigned types to say a number will never be negative. Instead, use assertions for this." [0]
[0]: http://google-styleguide.googlecode.com/svn/trunk/cppguide.h...
"Never use a signed type for a number that can never be negative"
One of my pet peeves is developers using int (instead of unsigned ints) for primary keys in database tables.
"SQL only specifies the integer types integer (or int), smallint, and bigint." http://www.postgresql.org/docs/9.3/static/datatype-numeric.h...
In this case, it makes absolutely no difference at all. It could be argued that writing unsigned int would make the code slightly harder to read. That said, I like to use stdint.h and unint32_t would, I think, not have any drawbacks.
> there is lot of code out there with interfaces expecting signed ints even though they should using uint
That's not a good reason to not use unsigned integers, it's a zero-overhead cast from unsigned to signed (at the risk of overflowing into the negative).
Using uint limits the optimizer...
Every time I see someone mention the optimization argument for signed integers I ask for examples and I've yet see a good one.
for (int i = x.size() - 1; i >= 0; i--) ...
for (int i = 0; i < x.size() - 1; i++) if (x[i] < x[i+1]) ...
Both will blow up badly with unsigned ints.(Well, to be fair, both will blow up with signed ints if x.size() is greater than 2G, so it's a matter of expectations.)
for( size_t i = x.size(); i-- > 0; ) ...Remember that a for loop does something like this behind the scenes:
size_t i = x.size();
while( i-- > 0 )
{
// your code which needs backwards iteration here
; // do nothing because there is no third statement
}Still, the trick makes it look suspect and that's an argument against using it.
This is true. The code is confusing to people not used to it. A workaround could be to hide this code inside a macro, so people not interested in digging into the code would take the macro's word:
#define REVERSE_LOOP( x, i ) for( size_t i = x.size(); i-- > 0; )
But unfortunately, that doesn't help with the fear that people has against unsigned types.That yields compiler warnings for signed vs unsigned then, no?
// count up
std::size_t i = 0;
while (i != 10)
{
std::cout << i << "\n";
++i;
}
// count down
std::size_t i = 10;
while (i != 0)
{
--i;
std::cout << i << "\n";
}
After initialization a for statement repeats "test; body; advance", this is ideal for counting up loops, but what we need for counting down loops is "test; advance; body". Since C/C++ do not provide the latter as a primitive you have to use a while loop as shown above. Using a signed integer to shoehorn a counting down loop into a for statement at the cost of 1/2 your range is a hack IMO. Note that when working with iterators you have to resort to a while statement as iterating past begin is UB.I'm pretty sure that it's just because "int" is one word and "unsigned int" is two, plus more than twice the characters. I suspect if "int" defaulted to "unsigned int" and you'd have to specify signed ints explicitly, the taboo would be reversed.
Never underestimate the power of trivial inconveniences.
There's example code on the CERT secure coding guidelines here (look under 'Substraction'):
https://www.securecoding.cert.org/confluence/display/seccode...
Writing safe code to calculate the absolute difference between two unsigned integers is much less hairy: max(x,y) - min(y,x).
Really, working directly in fixed-precision arithmetic is absurd. In order to be able to rely on its correctness with any degree of certainty, you need to very carefully track each operation and its bounds, at which point you may as well have just used arbitrary-precision types, explicitly encoded your constraints, and had the compiler optimize things down to scalar types when possible, warning when not.
Fixed-precision arithmetic has one main advantage over arbitrary-precision arithmetic: it is more time- and space-efficient. This advantage only applies if the fixed-precision arithmetic is actually correct and the fixed-precision arithmetic meets some concrete time or space constraint which arbitrary-precision arithmetic fails to meet. It generally takes time and effort to demonstrate that these conditions hold; because one can rely on the correctness of arbitrary-precision arithmetic without doing so, arbitrary-precision arithmetic should then generally be the default choice.
This assumes that you care about making relatively strong guarantees about the correctness of your programs. If for some reason you don't, then sure, use ints and whatnot for everything. If you do, though, I suspect you'll find that it's easier to track down a performance bottleneck caused by using bignums than an obscure bug triggered by GCC applying an inappropriate optimization based on overflow analysis.
if (index < 0) { /* error */ }
I die a little inside.Seems like a pretty ignorant pet peeve considering that's the only option for every database that doesn't auto-corrupt data.
Having a negative pkey space is actually useful. In LSMB we reserve all negative id's for test cases, which are guaranteed to roll back. This has a number of advantages including the ability to run a full test run on a production system without any possibility of leaving traces in the db.
1) If you don't need negative numbers, use unsigned integers.
2) If you don't need the extra positive range of unsigned integers (or defined wrapping), use signed.
You advocate (1), but C is generally based on (2), with the default int being signed, and many standard functions using plain int.
[0] though several do support UUIDs, which are essentially unsigned 128-bit ints, and which (with a well-selected generation mechanism) are better as server-assigned surrogate keys than sequential integers, signed or unsigned, anyway.
Furthermore, casting uintx_t to int and back again while using shared libraries is a huge pain in the ass and can waste a lot of programmer time that would be better spent elsewhere, especially when working with ints and uints together (casting errors, usually in the form of a misplaced parenthesis, are pretty small and can take a very long time to find).
2 billion survey results was never going to happen. 32,767 would have been fine as well except to compound the issue ops pointed the production site at the test database.
uintN_t (and intN_t) are MORE portable and cross platform than int in the sense that you get much better guarantees about it's size and layout.
Furthermore, int is NOT the size of the register (x64 commonly has an int of 32 bits) so any updating you'd have to do to uintN_t, you'd have to do to int as well. Regardless, I can't imagine why you'd need to do any updating in the first place - it's perfectly valid to stick a uint32_t in a 64 bit register.
> nevermind the fact that int is usually more optimized than uint these days
Where are ints more optimized than uint? Not in the processor, not in the compiler (modulo undefined behavior on overflow) and not in libraries.
However, I think this is a problem. The expected value ranges of your variables don't change just because your memory bus got wider - maybe you can use more than 4GB memory in a process now, but it's a mistake to plan for single array indexes being more than 32bit.
If you do try to be more flexible, I'm sure this would introduce more bugs than the forward-compatibility it'd add. Especially if 'int' is smaller than on the platform you tested on. That's why languages like Swift, Java, C# always have 32-bit int on every platform.
> casting errors, usually in the form of a misplaced parenthesis, are pretty small and can take a very long time to find
Agreed, but writing casts also adds unwarranted explicitness. What if someone made a typo and put the wrong type in the cast? How do you tell what's right? What if you change the type of the lvalue or the casted value? Now you have to think about each related cast you added.
What's the alternative? Well, the compiler should just know what you mean…
This is why we have uint_least8_t and friends. In fact, int is really just another int_least16_t.
> Furthermore, casting uintx_t to int and back again while using shared libraries is a huge pain in the ass and can waste a lot of programmer time that would be better spent elsewhere
Could you give an example? It sounds like you're just talking about performing the casts, which shouldn't take much effort at all as indiscriminately as C casts about integral values.
* dynamically test your program with ubsan to be sure they really don't happen, and then
* let your compiler optimize with the knowledge that integers won't overflow.
This last one eliminates maybe half the possible execution paths it can see, and loop structure optimizations practically don't work without it.
On the other hand, unsigned overflows? Some of those are bad, but some are fine, right? How will an analyzer know which is which?
Some notable libraries like C++ STL want you to write loops with unsigned math (size_t iterations), but those people invented C++, so why would you trust them with anything else?
If a function must-overflow the optimizer (hopefully) replaces the entire thing with an abort under ubsan, so you could look for that. But that's probably not sensitive enough.
And if the function is just 'x + 1' that may-overflow, but it's not important.
Maybe you want this: http://pdos.csail.mit.edu/papers/stack:sosp13.pdf
https://www.youtube.com/watch?v=9bZkp7q19f0
EDIT: I just realized that YouTube also posted a comment to that effect just below the video. :P
Being Korean might've given it crossover appeal into much of Asia? Just a guess.
I think they gave people who would otherwise have looked down on a silly craze an excuse to enjoy the video.
I'm sure some software will need to be re-written between now and 2038, but I don't think it will be quite as bad as Y2K just because that was only a 15 year gap (Sometimes less), whereas this is over 24 years.
I just think a lot of software will be naturally replaced between now and then. And while there will be a slight mad scramble to fix stuff at the last minute, I don't think it is Y2K-2.
See http://en.wikipedia.org/wiki/Year_2038_problem for a basic overview. None of this stuff is unfixable. But it is a real problem, and tracking it down will be hard.
I doubt this will be a real problem in 2038, then again the prevalence of computing devices is much larger now and will continue to grow by 2038, but so will technical aptitude, so hopefully they'll cancel out and this will still not be a problem.
As far as I know, the only major operating system that has dealt with Y2038 is OpenBSD [3].
[1] - https://github.com/opensource-apple/xnu/blob/bb7368935f659ad...
[2] - https://github.com/opensource-apple/xnu/blob/bb7368935f659ad...
[3] - http://www.openbsd.org/55.html , http://www.undeadly.org/cgi?action=article&sid=2013081307224...
Also, there are perfectly valid applications that require numbers of 8, 16, 32 or 64 bits (or variable encodings with arbitrary precision). Petabytes, embedded microcontrollers, etc.
But it's real?! It seems incredibly absurd that it could actually overflow, how are signed values useful for a count of views? How are you going to have negative views?
http://stackoverflow.com/a/1555186/67591
TL;DR: Google's C++ coding standard says: "Document that a variable is non-negative using assertions. Don't use an unsigned type."
As someone who has written a lot of C++ code to interact with hardware: I view an unsigned as a bucket of bits, and a signed as a number.
video_B.views > video_A.viewsGiven that switching to unsigned saves you exactly one bit, it's just not usually worth it. How often do you need exactly 32 bits of unsigned space, when 31 isn't enough, and you can't use 63? (I'm talking about standard-length integers here, not extreme situations where you're trying to make maximum use of 8 bits of storage or similar.)
Direct link to the video: https://www.youtube.com/watch?v=9bZkp7q19f0
Regardless, it was bound to happen sooner or later as youtube is getting older.
I remember when Twitter had rolled over their tweet ID's because they were using an int type that was too short. Should have gone with variable length strings to avoid that problem.
[0] https://en.wikipedia.org/wiki/Globally_unique_identifier#Alg...
ObjectId is a 12-byte BSON type, constructed using:
a 4-byte value representing the seconds since the Unix epoch,
a 3-byte machine identifier,
a 2-byte process id, and
a 3-byte counter, starting with a random value.
You can actually convert the _id to/from a timestamp, which lets you do cool things like never keep a timestamp field (you can convert a datetime to an ObjectId and use that for comparison).12 bytes is enough to just store a random value, with low chance of collisions until you've got around 100 trillion items. It confuses me why anyone would want to waste bytes of an ID on low entropy values like machine and process ID.
Or any combination of high res time plus random works nicely.
OTOH, Mongo's not exactly been a bastion of engineering excellence.
It isn't a uint32, it is an int32.
Using strings avoids one problem but introduce a bunch of others (e.g. a string is harder to verify, therefore less secure, and therefore needs to be handled with the kiddie gloves). Checking that every character is between 0-9 and dropping all other characters is easy, cheap, and effective. Then just check it is between uint64.Min and uint64.Max, and you're done.
uint32 gives them twice as much capacity (which isn't enough at this stage), they'll likely want to go with a uint64.
And you're doing this... how, exactly?