Clockwork raises $21M to keep server clocks in sync
techcrunch.com
techcrunch.com
* On LAN: NTP can get up to 1ms accuracy and on WAN about 10ms accuracy.
* People's 'system clocks' use NTP for synchronization but can still be completely off.
* When you call DataTime.now() in Javascript it returns the unix timestamp in UTC time.
* Since it is UTC it will have the same value anywhere in the world - however since it is set based on the the system clock it isn't guaranteed to be very accurate at all.
* Browsers now have a high-precision API for doing measurements of elapsed time called 'performance.' Don't use DateTime.now() for this.
* HTTP servers seem to return a Date field that has a unix timestamp in it. Some projects have attempted to use this to synchronize time in Javascript.
* There appears to be no good, robust Javascript libraries that can synchronize and keep time accurately. Some libraries exist that use a single server + NTP to try calculate clock drift, however. But this isn't as good as the NTP daemon.
* Researchers HAVE been able to write software to synchronize a clock with better accuracy than NTP using distributed networks -- this should be closer to what a lot of people are interested in. Here's a relevant paper I found on this: https://scholar.google.com/citations?view_op=view_citation&c...
* People have done some pretty cool bench-marking hacks to measure elapsed time in Javascript before the existence of the performance counter API. With performance.now() -- its not a clock but a counter and browsers can intentionally limit its accuracy to make 'fingerprinting' harder. https://stackoverflow.com/questions/6233927/microsecond-timi...
The primary reason that performance.now()'s precision is limited is for security reasons.
Spectre showed that it was possible to perform timing attacks to leak other memory on the system, and having a browser leak memory from other processes is quite a dangerous attack. As part of spectre's mitigations, browsers began limiting performance.now()'s precision.
See docs on this: https://developer.mozilla.org/en-US/docs/Web/API/Performance...
Chromium's impl: https://chromium-review.googlesource.com/c/chromium/src/+/85...
Total aside, but it's also more accurate to say that performance.now() is a monotonic clock rather than a counter. Counters don't necessarily have relationships with elapsed time, but performance.now() does, and it's conceptually the same as 'CLOCK_MONOTONIC'.
The problem the OP company is trying to solve is quite a bit tougher.
Of course it sounds extremely stupid, which is why there are so many drive-by "NTP solved this already" comments from people who don't deal with millisecond precision at high accuracy requirements, god forbid sub-millisecond.
> it returns the unix timestamp in UTC time
Unix timestamp is ALWAYS the number of seconds since January 1st, 1970 00:00:00 UTC. It cannot be in any timezone.
The difference arises from the fact that in utc some days are 86401 seconds long, but unix timestamp just repeats the last second of the day instead of being simple ever increasing counter.
As an example, since an ancestor comment was talking about JavaScript:
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... says Date.now() “returns the number of milliseconds elapsed since January 1, 1970 00:00:00 UTC”, which wording would suggest the inclusion of leap seconds (since they have certainly elapsed).
However, https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... mentions that Date.now(), Date.parse() and Date.UTC() ignore leap seconds; what it fails to mention is that actually it’s just that ECMAScript follows POSIX for time measurement throughout, though in a milliseconds base rather than seconds, meaning that everything to do with Date ignores leap seconds. (Citation: ECMA-262 §21.4.1.1 <https://262.ecma-international.org/12.0/#sec-time-values-and...>; warning, 6.7MB HTML document, slow to load.)
https://www.eecis.udel.edu/~mills/ptp.html
PTP takes advantage of special Ethernet switches and other devices which can decode or manipulate the time tags in hardware because yes your switches add latency.
This just sounds like more horseshit out of SV. I applaud them for convincing someone to give them money for this snake oil.
> Those people [fintech, etc.] are interested in our stuff, so obviously [that] means that thing [PTP] is not perfect for them. Part of the problem is, you’re trying to measure individual packets that are going through and get times. It means you have some clock there that you’re reading, and when it goes into the switch, and when it comes out, like, what is that clock doing in terms of its varying in frequency, and stuff like that. And then you get one sample and try to make things of it. It sorta pales in comparison to this big data approach we’re looking about.
This seems like a mischaracterization of PTP (802.1AS), which does use multiple samples to syntonize.
[1] https://www.clockwork.io/resource/ [2] https://www.youtube.com/watch?v=Opf9CBwP5R8
Would like to understand if the founders consider their approach to be a viable alternative to "atomic clock mode", but without actual atomic clocks.
They use GPS for global sync and a local atomic clock as reference. From conversations over the last few years this seems to be the common setup at big DCs, as the cost for a couple atomic clocks is minuscule compared to the scale of everything else.
No, Spanner has atomic clocks running in each data center.
And even in HFT, it is silly to try to use your x64 system time to get nanosecond precision. Even if you are really good, using bare hardware, all your code and data is in L2, an x64 system would still be giving you 100ns level jitter. No matter what you do. What you want is to have your network card (or FPGA) timestamping the packets accurately. And then releasing these packets at a precise timestamp.
Saying that, maybe Clockwork is doing something completely different. And not in nanosecond level, but at a millisecond level. They talk about "measure true one-way delays of a packet or remote procedure call, discover network bottlenecks and "hiccups" (outages lasting a few seconds), and identify underperforming VMs arising from "noisy neighbors." Now, that could indeed be useful.
I wonder how this deals with the numerous retimers and DSPs required for really high speed ethernet.
The GPS locked master clocks are custom hardware (ex: https://evertz.com/products/5700MSC-IP), but a lot of the edge devices like the video playout servers are standard X86 hardware with Ubuntu Linux, with PTP delivered in-band over the 10G/25G ports to the device (same link as the video/audio flows).
(… Which I realize is not as accurate as the newer White Rabbit related PTP update you refer to - but still odd that the original article referenced NTP and not even regular PTP which has been in use for a while and seems pretty close to their claims even before the more recent enhancements)
It's all fairly simple in the end if the hw has the tools but good luck buying SyncE or 1588 without $$$.
In hardware that's actually designed to do this, I think it's called "synchronous ethernet" but you can totally duct tape it into pedestrian hardware the way I did.
It does need L2-L1 coordination but the ethernet MAC usually has no idea of what's happening on the L1 (it speaks "medium independent interface" or *MII). The CERN people warn against using Base-T since the PHY needs to do pretty complex signal processing which would likely destroy the syntonization.
They then use basic fiber links, but 100g, 400g+ have a gearbox and DSP and FEC in the way.
By using a high volume of requests, they can actually average out a lot of the well-behaved jitter sources.
Hence my initial question: are they adding extra information to actually achieve those stated goals, or are there algorithms just "probably" better, and in the latter case what are the use cases? Distributed transactions and whatnot are fundamentally broken if your "better" synchronization might still be wrong.
The graph cycles are neat, but they even admit in their own paper [0] that the approach is limited to half the max path asymmetry (RTT/2 if the asymmetry is totally unknown) -- pure software clocks have hard lower bounds on accuracy that can't be overcome without additional information (which digging elsewhere on their site it looks like they do actually integrate with sources like GPS antennas).
The rest of it is actually pretty interesting; in a datacenter context you might very well have low asymmetry, and everything else seems well done and likely to be much better than NTP for common scenarios.
[0] https://www.usenix.org/system/files/conference/nsdi18/nsdi18...
it doesn't have to be, market will decide.
Because it turns out NTP can be much more accurate than most people realize.
From the Chrony FAQ[1]:
> When combined with local hardware timestamping, good network switches, and even shorter polling intervals, a sub-microsecond accuracy and stability of a few tens of nanoseconds might be possible
Good network switches and NICs with hardware timestamping support are commonplace now in server environments. NTP with Chrony is pretty hard to beat in terms of simplicity, reliability, and accuracy.
Those don't care about the protocol as they can provide hardware timestamps for all received packets.
Other popular NICs like the Intel X540 or XXV/XL710 are limited to timestamping of PTP event messages in order to limit the rate of timestamps which needs to be handled by the driver. For those chrony supports an NTP-over-PTP protocol which forces the hardware to trigger the timestamping by wrapping NTP messages in PTP.
Sub-microsecond accuracy is certainly possible. Here is an example with 3 network switches: https://chrony.tuxfamily.org/img/client-hwts-3switch-f323.pn...
The accuracy is limited by asymmetries in network switches.
In any case, whatever algorithms Clockwork is using with their protocol, I'm sure they could be used with NTP too. If additional information needs to be exchanged between the hosts, extension fields can be specified for that.
TrueTime, PTP, FB’s time card, all solve one problem that is to provide a bound where “true time” falls in. You need to be able to _guarantee_ the precision. If you say the true time lies between [T - d, T + d], it must be the case. The spanner paper provides data that the probability of TrueTime being wrong is less likely than random hardware failures (bit flip, etc.). Nothing is 100% in computer, but once you have like 20 9’s of reliability, our society has collectively accepted it as good enough (like we assume hash collision would never happen).
Now AFAIK, no machine learning model can come even close to 10 9’s of accuracy. Clock, to me, is a piece of foundational infrastructure that should provide a very solid and simple mental model so we can reason about it and build other things on top. I am skeptical that NTP + ML would work.
Wouldn't it be easier to just make distributed servers deal with large 'packets' or large individual tasks on their own?
If you don’t try hard you are likely to end up with something accurate to a few ms. AWS has some service which can get you synced to around 300 mics in normal conditions. But 300mics is pretty bad.
Doesn’t this company basically require their customers to not use cloud providers or have they figured out how to get good clock sync despite cloud provider networks? It seems limiting if they don’t work in the cloud.
I think the most common problem we have with clock sync (at least the most common problem I see) come from overloaded network cards slowing down timekeeping packets. I wonder if that’s much of a problem with this company’s solution.
I played a lot with an asynchronous chip, meaning it didn't have cores, it had computers, and they didn't have a frequency. Then I connected it to an oscillator (it's easy, just lay one wire) and was getting sub-nanosecond overall measurements with no jitter.
And it's like, why can't you hard-code that in? Like in your test suite, test the code and see if it's as fast as it's supposed to be, or so fast it's clearly not doing the work.
Synchronizing 5G Mobile Networks:
https://www.ciscopress.com/store/synchronizing-5g-mobile-net...
I’m curious, who is even the target customer here? Places like FB and Google are working on the problems themselves.
They talked about some “latency sensei” product or something but what is it gonna tell me? How many nanoseconds it took for my API request to go from a load balancer to the web server, with 5 ns precision? Is that precision really needed?
Does it matter that a different server picked up the message first?
Now what if the server just thinks it's eight feet to the left because the clock is out by several nanoseconds?
Can't you just ignore that, and still use the server that got the lowest timestamp?
What's the practical difference?
When the servers can communicate in x amount of time, the question was specifically about clock drift "much smaller than" x.
Who gave this guy who's apparently never heard of kerberos, ceph, SAML, or any other technology the rights to interview? Sure, their windows are larger, but "nobody uses time" is so clueless it blows the mind, and it's hard to imagine who decided that "NTP, but machine learning" is a $21M idea
If these folks can achieve what they want at a couple orders of magnitude cheaper than current prices then you are absolutely going to see a lot more regular use cases show up to take advantage of it.
No, not really. I work in finance (HFT, systematic market making) and this industry heavily relies on high precision clock sync. It's even regulated by law (MiFID II); in practice everyone on the street uses PTP because if you don't know how fast you're really going or where your latency spikes you already lost the race, just don't know about it.
There's a couple of other domains where micro- and nanosecond level time sync is of paramount importance. PTP has become much cheaper over the years especially when you are in a data centre and can get it as a service instead of setting up grand-masters, slaves, GPS antennas, etc.
This statement just reads like badly researched topic on their part.
I considered starting a competitor in 2018 when I first saw this company, but I don't have the same connections to customers that these guys do.
My conclusion was that PTP precision with software timestamps would be a good company, but not something between. I hope they can prove me wrong!
If one has such an accurate clock they might be able to accurately plot their position on earth to within 1 ft merely by checking the timed pulses of orbiting satellites.
Hey look we got a GPU to learn the Kalman filter algo.
Many are actually better than 10 ns basically a 1 foot CEP is 1 ns. If they are fixed and always on the accuracy it can build over a day is incredible.
If you arrive to the railway station 10 minutes +-2 minutes before your train leaves you are using time. In a very real sense. In robotics we sync multiple computers on the robot to about miliseconds, that is using time. In a real sense.
Maybe more accurate, cheaper sync will enable more applications. Maybe it is a good business to specialise on it. But saying that nobody excep for Google, CockroachDB or someone doing database things uses time is bulshit and can be called out for what it is.
Why are you in a very specific technical discussion correcting what is obvious? Yes, you're using time. You're not using time at the precision this company is targeting. At that precision, there are few current use cases, but if they make it cheaper, there will be more. What's there to argue here? You're offended you're not being included as users of time?
"We're lunching more accurate satellite imaged maps. The current users are largely nation states and niche industries due to its cost. This will enable more map use cases"
I use Google Maps all the time. I used it for my last trip! How dare you.
Please refrain from personal attacks. It is unecessary and doesn’t add to the conversation. Thank you.
> I use Google Maps all the time. I used it for my last trip! How dare you.
Do note that your example didn’t say “nobody uses maps”. If an imaginary salesperson would say “nobody uses maps” I would call bulshit on that too.
> Yes, you're using time.
Great, so the company representative should simply not say “nobody is using time”. And the commenter I was responding to shouldn’t say that applications which are fine with a lower accuracy “aren't ‘using time’ in any real sense”.
Heck, all I know maybe nanosecond accuracy is where the bees knee is, and everyone is going to be amazed by the awesome new applications it is going to enable. But you don’t need to disregard all the history of timekeeping and all the current applications to make that point.
On 28 April 1789 a mutiny broke out on HMS Bounty. The ship’s former captain, departing on a rowboat, demanded that the mutineers give him K2 the ship’s chronometer. The mutineers refused, because the chronometer was worth about as much as the whole ship, and they needed it for their onward navigation. Somebody should have told them that they are not using time in any real sense. Probably they would have laughed at that idea.
I don't see how something like Kerberos or TOTP, that fails without some sort of synchonization, can be defined as not "using time" in a real sense. (If cluster nodes were drifting by seconds, I'd be checking them, even if that was only going to affect something like make.)
For GPS to work at all, you need very accurate knowledge of the offset between each satellite's signal.
The only way I can think of to have precision but not accuracy is if you design a system where there's a buffer between the antenna and the signal processing and you don't know how long the buffer is. Is a design like that a practical concern?
The figures (eg $21M) and names dropped (eg Stanford) are an appeal to authority, which does make me curious.
I’d love to see some papers, I went through all of Balaji Prabhakar‘s publications (titles only), and didn’t see a single paper that sounded like “NTP but machine learning.
If anyone else knows more about this, I’d love to hear from you. Surely $21M doesn’t get dropped with at least someone doing due dill on the tech?
Clock synchronization: https://www.usenix.org/conference/nsdi18/presentation/geng
I think the TechCrunch article doesn't really explain their application of clock synchronization well. Here are the other relevant papers and my attempt at explaining the general idea below.
Network edge-based measurement (one-way delay?): https://www.usenix.org/system/files/nsdi19-geng.pdf
Congestion Control: https://www.usenix.org/system/files/nsdi21-liu.pdf
Each paper I've listed builds on the last. Their method to synchronize clocks made accurate and efficient measurement of one-way delay possible with commodity hardware. This measurement of one-way delay occurs at the edge, allowing them to "hold" incoming packets at the edge for extremely small periods to reduce congestion (latency) while maintaining throughput. From my understanding, traditional congestion control algorithms require rich telemetry from the entire network, which is likely not accessible in a public cloud environment. Balaji and clockwork's algorithms only need to make these measurements from the edge (which customers in public cloud have access to).
I'm curious to see how all this will scale for multi-region deployments. If the latency between VMs from region 1 and region 2 is significant, I wonder if the measurement will actually be useful in deciding to "hold" the packets.
Accurate clocks sync enables true one-way delay measurements (instead of RTT/2), this allows for edge-based network visibility. We launched Latency Sensei beta – a sensor, monitor and auditor that provides visibility into cloud deployments. The gallery has cloud fitness reports on GCP, AWS and Azure. Some interesting reports include: 1)how VM colocation impairs network bandwidth, and 2) tale of 2 cloud regions, London vs Singapore. Take a look and we'd love to get some feedback https://sensei.clockwork.io/user/gallery/
On Congestion control, an edge-based solution is coming soon. If you're interested in a private beta, email us at hello@clockwork.io
The first article you reference includes something about Support Vector Machines (SVMs).
They have the 'names' because they are a reputable group of people. They have been at this problem for a while, first with extensive research that required clock sync as a pre-requisite, and then to evolving into clock sync as a formidable problem unto itself.
Clockwork is a rename of the company as far as I can tell. Their original name (Tick Tock Networks) [1] was probably too close to what has become a very popular homophone.
You seem to know their work, so if you have further publications, I’d love to get them please.
Haven't watched this one, but might be complementary: https://www.youtube.com/watch?v=Opf9CBwP5R8.
Juicero raised $120M.
Just some words:
Theranos.
Ubeam.
Moller.
And many, many others. Investors do this all the time. FOMO + 'wouldn't it be nice' are powerful ways to part fools from their money.
You’d be surprised, I was.
It was fascinating to get to watch the process. I'd always taken it for granted the big hotels, railway stations, and churches would often have big clocks you could see from a distance and that they'd be self-correcting. But of course a lot of old clocks are purely mechanical so when the time changed an hour forwards/backwards somebody would need to physically change them.
Bless you. Are you younger than 40? :-)
Tech people who are over about 40 have, alas, no trouble imagining this at all!
I've spent a _long time_ working in distributed systems with components which are "close to the metal". There are use cases for this, many of which are better solved by "hardware timestamping NICs" and "run a stratum 1/2 NTP server with high locality (and if you're geo-distributed, run multiple on systems dedicated to that purpose, because ensuring you're within nanoseconds of _someone else's_ infrastructure is generally an order or magnitude less important than internal coherence).
That was the best possible quote from the author/subject trying to explain the use cases. Did you even read TFA? It's completely unclear from their article what the intended uses cases are, what the problem space is, and how much working knowledge they have of existing solutions.
Naming people who invested in other companies who are investing in this one and showing what looks like a dashboard for a timekeeping solution is normal TechCrunch garbage, but that quote was a step above.
Comments like yours (going all the way back to /.) days really just tell me that you didn't read the article, and it makes it look like your primary goal is sophistry.