Interestingly enough, even the regulators don't have good (only millisecond-resolution) trade data.
Interestingly enough, even the regulators don't have good (only millisecond-resolution) trade data.
The trading before the official announcement is a concrete proof, if this is official, and bug/error free. But then again, this is not my forte.
More importantly, why wouldn't you invest in colocations and collect the data yourself using direct feeds (with GPS clock synchronization etc to validate the data)?
please send me a line, when you can. Cheers.
So you look at yesterday's view count and today's view count and notice that somehow the view count only increased by 900. Something is afoot!
Nanex's argument is tantamount to saying "that means we must have lost 100 organic views today"
My argument is tantamount to saying "I actually bothered to look at our access logs (which we record on our servers) and only saw 900 that we could definitively attribute to real Facebook users. Is it possible that the report is incorrect or falsified?"
Now I bring up this example because this actually did happen with facebook. Quick HN search revealed one such discussion: http://news.ycombinator.com/item?id=667308
Back to the current situation. There are many sources of market data. Each individual exchange generates its own feed, and with the major exchanges (NYSE, NASDAQ, BATS, DirectEDGE, ...) you can colocate in the exchange data centers (NYSE and ARCA are in Mahwah NJ, NASDAQ is in Carteret NJ, and various other exchanges are located in New Jersey and Chicago) and record the data yourself. There is a unified tape (CQS/CTS) which combines and disseminates a combined record (across all exchanges). This is used to determine the "national best bid/offer" -- the prices people are willing to buy/sell at.
The process of CQS generation is fraught with problems, but lots of older traders and academic types use CQS data because its much cheaper to get that data than to get data from individual exchanges directly. However, you are subject to the quirks of the combination process, including subtleties regarding timestamping data (since this data includes trades and quotes from Chicago and from New Jersey, the sequence of events may appear different if you record from chicago or new york or philadelphia or some other place; if you ask the exchanges to timestamp directly, you have to worry about clock delay and skew between the exchanges' servers).
nanex is saying that it is acceptable to depend on that data and any anomalies must have occurred outside of the recording process. I am saying that the recording process can create the types of anomalies that nanex is showing, and that the only way to be sure is to record the data directly and carefully synchronize your recording machines. AND when you do that you see that there really is no anomaly.
Just to emphasize how sloppy the exchanges are with regards to timing: on the BATS exchange they use multiple servers to run trades and generate quotes, and every once in a while you see messages appear to be out of time order because the individual machines weren't properly synchronized (although, if you filter for a single ticker, messages are always in chronological order)
They should just spring for the data before making accusations -- the problem is that when you cry wolf all the time no one will take them seriously when a real case comes around.
Suppose you left ntpd running and automatically adjusting the clock every hour.
If your clock is running faster than pool.ntp.org, and you are synchronizing to it, you may end up adjusting in the middle of an event. Because your clock is running fast, you would jump back in time, breaking the sequence of time (this is somewhat equivalent to what you see during daylight savings time if you aren't intelligent in the way you handle the backwards hour shift)
In this case, if the adjustment was forward in time, there would be a gap.
However I can't remember where I read that and could be totally wrong.
For those who do latency tests, this is a very important point: you should always be on the lookout for what clock is recording the 'start' and the 'stop' and to be sure to consider clock skew.
To get a sense for how far timestamps can diverge, OATS -- the reports that are sent to the Financial Industry Regulatory Authority -- require that machines be synced to within 3 seconds of NIST (which is nearly 7.5x longer than the 400ms quoted).
When you measure latency, you always have to be careful about clock issues at the point where you measure the start and the point where you measure the end. This is old-hat for sysadmins and others who deal with these types of issues. This is why round-trip latency numbers are easier to work with: both the start and the end times are taken on the same clock.
It's clear, given nanex's responses, that they depended on someone else to give the timestamps. Before concluding that someone had inside information ahead of time, they should check their processes. It's like someone claiming they built a perpetual energy machine because they confused power with energy (i want to say it was paul newman but the name escapes me -- this actually happened)
And I can't see any particularly efficient way to level the playing field with a government that barely understands facebook and fair use. Just try to explain microsecond latency to them, go ahead, really, should be even more fun than a series of tubes(tm).
The quotes themselves are provided with resolution from 1msec to 1nsec (depending on the exchange and on the type of data you wish to receive)