Date parsing performance on iOS (NSDateformatter vs sqlite)
vombat.tumblr.com
vombat.tumblr.com
I did the same with Java's Integer.parseInt(...) method. It is an interesting task to go through.
Now I'll spend the rest of this rainy afternoon playing around with writing a fast ISO date parser :)
Edit: Seems Java's Joda Time library already does parse ISO dates really quickly. 7 seconds for 4 million on my MBP
Edit: A fast custom date parser for ISO dates I just wrote can parse 4 million dates in 150 milliseconds.
Although I'm not sure how many people on iOS need millions of dates parsed.
Wrong. That's very fortunate. Unix time stamps have some serious deficiencies as data type for storing time information: for one, they lack precision. One second just might not do it. Then they lack any time zone information. You will never know what a specific time stamp is in. GMT? UTC? Time zone where the server is in?
Sure. Maybe you are lucky and it's documented (it probably isn't because people who care about such things are not using unix time stamps to begin with), but using a string time stamp formatted in ISO means that no documentation is needed. The encoding is good enough to store any sub second time stamp including time zone info.
That way, you can turn any of these into whatever your environment uses internally which you will then use in conjunction with the library routines to deal with all the difficulties related to doing math with dates (how many days in a month? What about leap years? What about time zones? Not really hard issues, but many to keep in mind and many possible causes for bugs)
$ date -d @1378585039
Sat Sep 7 20:17:19 UTC 2013However, you're overall point there remains valid, because people will try to pass off something as a "UNIX time stamp" that is actually in a different time zone. There is value to self-describing data.
{timestamp: xyz, tz: "America/Los_Angeles"}
The ISO date format's notion of timezones is a compromise that gives you the worst of all worlds. They complicate referring to a physical instant, because you can refer to it in several timezones, rather than the unique representation of a unix timestamp. But they're inadequate for political time, because what time comes 6 months after 13:00 (+00:00)? (It could be 13:00 (+01:00) or 13:00 (+00:00) or likely others - you need a symbolic timezone like "Europe/Lisbon").
For physical or "system" times unix time is great (unless you need the greater precision, but how common is that?). For user-facing times ISO is inadequate. The use case for ISO string datetime formats is very narrow.
I thought it's supposed to be an external format. I'd always expect a computer system presenting the output information to the user in his local time zone, while accepting inputs from all time zones equally.
"But they're inadequate for political time, because what time comes 6 months after 13:00 (+00:00)? (It could be 13:00 (+01:00) or 13:00 (+00:00) or likely others - you need a symbolic timezone like "Europe/Lisbon")."
You can't standardize a changing practice. I'd never expect it to deal with these issues.
We're talking about a web API here, not user display. But even so, IME users don't think of their timezones as "+8" or the like, so for human I/O you want to use symbolic timezone names, not offsets.
Quote:
> Unix time, or POSIX time, is a system for describing instants in time, defined as the number of seconds that have elapsed since 00:00:00 Coordinated Universal Time (UTC), Thursday, 1 January 1970
That said, the exception for local time, at least in my opinion, is agreeing on dates in the future meant for human interaction (e.g. "I'll meet you at 7 AM local time in Time Square on the 3rd of April 2068"). Here time zone rules may actually change before the date transpires, and you can't be sure of the representation in any other zone or format until closer to the event.
Local time is a weird thing and changes all the time.
For giggles, look at the history of timezone rule changes in tzdata.
Most timezones have at least one duplicate hour per year (IE the same time occurs twice) in the US as well.
Local times are not an appropriate way to store time.
Note: ISO8601 does not give you local time anyway, since you cannot infer the timezone from the time offset.
That you cannot infer the timezone from the time offset in ISO 8601 is a good point though.
Or at the very least, preserving the time offset.
The advantage of ignoring leap-seconds on the recorder is you can map any sufficiently precise monotonic clock to UNIX time with a simple linear equation. Personally I think it makes a lot more sense to keep the complexity contained to the decoder, rather than the encoder where bugs could mean you end up not recording an accurate timestamp to begin with.
In any event POSIX time stamps are fine w.r.t. leap seconds, it's the conversion functions which may or may not reflect them.
As others have mentioned, Unix Timestamps can be arbitrarily precise by adding arbitrarily many places of decimal precision (and this is common practice, supported by the Unix "date" command, among other things).
Secondly, Unix Time is an absolute timescale that is not relative to any time zone. A Unix Timestamp alone unambiguously (1) identifies an absolute point in time; there is no need to involve time zones, which are a political concept. A Unix Timestamp can be converted to any timezone and vice-versa. Any representation that is based on civil time is going to be more complicated and have more edge cases.
Thirdly, time zone offsets like -03:00 do not actually specify a time zone; they specify a time zone offset. These two are not the same thing. There are multiple time zones that can have a -03:00 offset, depending on the time of year. Even given a specific time of year, the time zone offset may not uniquely identify the time zone. For example, Arizona doesn't do daylight savings, so if you see a -07:00 time in the summer it could either be a PDT time (used on the west coast) or a MST time (used in Arizona).
Unix Timestamps have many advantages over text-based timestamp representations. They are much simpler to parse and have far fewer lexical variations. They are never invalid (whereas text-based dates like 2000-01-32 can be). They can be stored directly in a numeric variable. You can perform math on them directly.
(1) Except for leap seconds
The same hour does also not occur twice in unix timestamps, though it does in most timezones (but not time offsets). Conversion rules are a mess, and have changed over time.
[1] https://code.google.com/p/crush-tools/wiki/ConvdateUserDocs
Shelling out for batch-processing loads of data is on the other hand great.
strptime_l took 58.803 seconds
NSDateFormatter took 107.570 seconds
sqlite3 took 7.022 seconds
And with MishraAnurag's suggestion of using timegm instead of mktime:
strptime_l took 21.656 seconds
NSDateFormatter took 108.163 seconds
sqlite3 took 7.096 seconds
My c is quite poor so if you have any suggestions on how to improve I'd love to hear them!
Two different design goals. Two different sets of design trade offs. (For "design" in the Rich Hickey sense.)
The NSDateFormatter is already being cached. That was my first suspicion on finding this issue too. We are using one formatter per thread in the production code, but that doesn't apply for the code I've posted since everything is done on the main thread using a single formatter instance.
Yes, NSDateFormatter is slower than other methods including some C libraries out there or this novel approach for turning a string into a NSDate however in most instances it's plenty fast enough and has a bunch of useful functionality [1] the least interesting of which is easily turning a string to a NSDate.
If you are optimizing this aspect of you code first you are likely wasting your time and would suggest iOS/Mac developers get to know NSDateFormatter intimately especially if you are displaying date/time information to users anywhere in you apps.
You should wrap this piece of code in a nice API and let people benefit from your findings. (Someone else could do it too, I'm just saying!)
[Edit: I see now that in the test there is only 1 statement object ever created for a test of a million dates. Better than I thought initially. But my guess is the statement object still creates some degree of inefficiency not found in directly calling the C version.]
7 seconds is a long time in CPU terms, I am sure that he can do better.
But you are right. In production, it makes more sense to let SQLite handle the conversion and insertion in the same statement.