Twitter IDs to roll past 53 bits in a couple days (may break javascript apps)
groups.google.com
groups.google.com
You can use 64bit floats (double) to store exact integers up to 53 bits.
This works in a lot of languages on non 64bit hardware, php for example.
http://www.ietf.org/rfc/rfc4122.txt
uuids are made just for this purpose, string based, never bigger than 40 characters (with dashes and curly braces).
Most products use uuids or Microsoft's name for them guids.
I didn't know that Javascript couldn't handle numbers bigger than 53-bits, but honestly, these should have been strings from the beginning.
The JavaScript Number type can't handle more precision than 53 bits. Magnitude is orthogonal due to floating-point representation. Precision is governed by the size of the mantissa, which is 52+1 bits long in the 64-bit IEEE 754 representation used by JavaScript.
IMO, it's not a bad idea on their part. A 32 bit UNIX timestamp * 2 ^ 32 + a 32 bit sequential id let's them track up to 4.2 billion tweets a second and should work just find up to the year 2106.
Edit: As to why it's a good idea, you can have different systems handing out ID's without stepping on each other’s toes or even talking to each other. The full ID is composed of a timestamp, a worker number, and a sequence number. Granted, I would probably put the sequence number ahead of the worker number so sub second tweets are better ordered vs. being ordered strictly based on the system that generated them.
And once you've done that, why not just go all out and use UUIDs?
I don't think they can be blamed for using a trivial incremental key when they had 10 users, I am sure they were not expecting to have 200M :)
This seems to be the relevant id generating code: https://github.com/twitter/snowflake/blob/master/src/main/sc...
This is of course doable with strings too, once you decide what the ordering is.
In a more practical application, a timeline comes back and the top-most and bottom-most IDs are stored. When a user gets to the bottom an API call is made to load more so you look at the bottom-most ID. If they want to load new tweets you look at the top-most ID. No math needed, just looking up values. They could've been strings all along and it wouldn't have mattered much.
I don't understand how saying you're gonna access the timeline without looking at the values denies that.
Moreover, I did not object to using strings, I said you only want them to be orderable.
Also, it seems strange that they would include the new ID in string AND integer form in their JSON. I realize that they don't want to break existing javascript apps, but isn't there a significant bandwidth cost in adding that sort of kludge to the API when you're serving a quadrillion of these api requests every day?
In the new system, they don't have to coordinate every server just to make sure IDs are sequential. They just use the time stamp and some machine-specific information.
Adding two positive numbers should never result in a negative number. Adding a one to an integer should result in the next largest integer.
There are new languages created all the time that don't have built-in support for arbitrary precision integers or rational numbers. They might have some neat ideas, but if your language can't even get arithmetic right, it's garbage.
That's only because everything is an instance in Javascript. There are no classes, it's prototype OO.
As I have suggested in a post 9 months ago, and as I would have designed this system at any date since 2002 or so, the twit ID would be composed of userid bits + time bits. In the post below I suggested 32+32, but other divisions are acceptable depending on your "bot user" policy. Such an ID would at the same time be sufficient (up to 4G users, up to 5 tweets/sec/user AVERAGE, up until 2030). You can have two times as many users for just "2 tweets/sec average". Facebook only has 500M, so 4G should be sufficient.
Such a construct makes the entire system significantly simpler and more robust to "meaningful" failure. I haven't seen a single thing tweeter has done right in the technical sense.
They do deserve marketing and bizdev credit.
http://www.reddit.com/r/programming/comments/b2u6t/twitter_o...
You would be changing behavior in a way that breaks apps. Twitter doesn't want to do that.
Look at e.g. YouTube and many other sites around the same time. They knew what they were doing; Twitter didn't.
I _have_ actually designed such a system in 1999, that used 48 bits, and it worked perfectly well. (Only had 28 bits for the user-id, which would have been broken at the 250M users -- alas the system never had more than than 5M; This was in the years 1999-2003).
The only way you can shard absolutely monotonic is (effectively) randomly, which is an option however you assign ids; but other assignments let you build a much cheaper, much more robust system.
Twitter is an API. Even if they have the knowledge of how to fix past mistakes, they need to ask for feedback, give plenty of notice, and set a deadline of when the old version is cut off.
Maybe they didn't do everything perfect day one, but I don't think there are any APIs the size of Twitter. Cut them some slack, they're not morons.
There do exist independent youtube clients (though not as many as Twitters's), but using the encoding they did, youtube has made it so that it is never going to be an issue, whereas for twitter it has already been a significant issue twice (that I'm aware of).
It's very easy to dismiss sound engineering in retrospect as luck or as "how could anyone have known".