Why is my NTP server costing $500 per year? Part 1 (2014)
blog.pivotal.io
blog.pivotal.io
IIRC, the university called Netgear out for doing something stupid and disruptive, and Netgear stopped doing it. The second best possible scenario, I guess.
https://en.wikipedia.org/wiki/NTP_server_misuse_and_abuse#NE...
> NETGEAR has donated $375,000 to the University of Wisconsin–Madison's Division of Information Technology for their help in identifying the flaw.
Although that appears to be uncited.
Netgear issued patches for the devices. Most people never update their server firmware, and we're talking about over 700,000 devices. The university still gets considerable traffic.
https://en.wikipedia.org/wiki/NTP_server_misuse_and_abuse#NE...
For what you want to have any effect they'll have to sinkhole/throttle the traffic upstream before it ever reaches them and as a university they are effectively an ISP so that might not even be really possible.
I do recommend reading the incident report posted above if you have an interest in network operations, it's quite interesting!
It'd be interested in seeing if there's been any update since 2003.
E.g., is it really "considerable traffic" by 2016 standards? The original flood in 2003 was 150 MBps - I don't think I'd notice if I got a flood of 150 MBps on my home connection.
How many of those devices are still around 13 years later?
They might give you a web interface where you can configure certain settings (e.g. integrated Wi-Fi) but the ISP ultimately has at least some control over any cable modem connected to it.
And to be able to remotely change the code running a HUGE security issue.
If you're an admin who cares about their infrastructure, you're not using a bargain-basement Netgear router, and if you are, you'll have gone through every single menu and seen the auto-update option.
So, if you were the only server in the pool, perhaps you would get a lot of Puerto Rican traffic?
> I wonder if Puerto Rico has run out of its pool of IPv4 addresses. After Europe and Asia, just this month Latin America as well, have exhausted their IPv4 pools, many local ISPs have resorted to using NAT to deal with the scarcity of addresses (of course, after years procrastinating IPv6 and pretending that this day wouldn't come about). Given that the source is a Puerto Rican ISP, and one of the offending addresses from a small /21 network, it's possible that NAT is to blame. As ISP NAT increasingly becomes more prevalent, this is going to be rather touchy to deal with abuses. For is it an abuser or just several innocent users behind a NAT?
My bad. Now that the post has been picked up on HN, I'll try to write up a conclusion.
With the advent of the cloud I see a lot of businesses forgoing any central IdM solution like AD or IPA and as a result don't tend to have an NTP server unless they explicitly configured one.
All of my servers are joined to an AD or IPA domain, so they all use my local NTP servers by default.
In addition, even with minpoll, that's still a hell of a lot of clients abruptly polling him abusively.
> I wish he'd explained somewhere how they leapt to examining virtualized NTP clients...
I had a hunch [wrongly] that the traffic was caused by a particular operating system. I didn't have enough machines to run the tests on bare-metal, so I virtualized them. And I suspected that virtualization would provide a worst-case scenario (the virtualized clocks would be jittery).
My big surprise was that Windows was a model client (once per day), OS X was good, too, and that FreeBSD was the worst and Ubuntu a close second. It was a complete inversion of what I had expected to find.
> what they ultimately did (since there's no part 3 that I can find).
I never wrote part 3, but now that it seems the post has gained traction on Hacker News I might be inspired to write one. The summary would be as follows:
To reduce your costs by a quarter, use the following lines in your NTP configuration file to throttle overly-aggressive clients (the important directive is `limited`:
``` restrict default limited kod nomodify notrap nopeer discard minimum 0 ```
Here is a description of NTP rate-limiting and why `limited`, `kod`, and `discard` are important: https://www.eecis.udel.edu/~mills/ntp/html/rate.html
Here is my current NTP configuration: https://github.com/cunnie/deployments/blob/95e9c71e882d453ec...
Your configuration is lacking a number two big best practices. The most glaring is that you really need to add 'iburst' to your server stanzas. After that you should think about adjust minsane and minclock.
Enable KOD, by all means, but you may also consider putting in some (high) per-IP rate limiting for 123/UDP in your firewall rules as a backup plan (for if/when clients ignore kod).
My then company wanted to give back ny doing this many years ago and it was an eye opening experience. We had troubles almost immediately with utilization and script kiddies. The company ended up only doing it for a relatively short period and ended up making contributions to projects instead
Thanks, Spooky23. As you pointed out, contributing to the community takes more time than originally expected.
... unless they saturated the available bandwidth but, really, that's a different issue (although also preventable!).
What GPS module did you go with and is it still available? Did you have problems getting signal inside (need to be by a window, run an antenna, etc)?
I put the GPS antenna on my window sill, it has no problem at all staying locked. My plan had been to stick it outside the window but turned out not to be necessary.
I haven't tried it myself but I've heard of several other good experiences w/ the BeagleBone Black. The Garmin seemed to work the best for me, although it is a little more expensive. I was strongly considering putting a few of them in $work's (private) facilities as a fun, nerdy project but I never got around to actually doing it. The Garmin with a BBB might very well be a great combination for that.
One other thing: make sure you use a "real" serial (or parallel) port -- not a USB to serial adapter!
Digital Ocean is a great deal! Thanks for pointing that out.
The reason I use {aws,azure,google} to host my NTP servers is that my day job is developing a VM orchestrator (BOSH) for Cloud Foundry, and BOSH doesn't support Digital Ocean yet (AFAIK). But that's a personal choice, and an admittedly expensive one.
It'd be nice if BOSH could have detected this problem and warned us about it, ideally during deployment. But it'd be even nicer if we didn't suck at configuring AWS. If you could fix either of those things, that would be great!
Yeah, maybe that's why my AWS instance is the most jittery of my 4 timeservers (Google, Hetzner, Azure).
Pivotal bridges the Silicon Valley state of mind, modern approach and infrastructure with your organization’s core expertise and values. Who we are and what we do together can reshape the world
Your article suggests otherwise.
Edit: s/your/you're/
NTP runs fairly decently in a VM. Don't take my word for it — look at the graphs of my servers:
Here's my Google VM, notice the jitter is within +/- 5 milliseconds:
http://www.pool.ntp.org/scores/104.155.144.4
Here's my Hetzner VM (Germany). +/- 10 milliseconds, though I can't help but suspect the distance from the monitoring station (Los Angeles) may have more to do with it than being a VM:
http://www.pool.ntp.org/scores/78.46.204.247
Here's my AWS VM. Much worse than Google in that it's +/- 50 milliseconds, but still good enough to pass muster with pool.ntp.org:
http://www.pool.ntp.org/scores/52.0.56.137
Here's my Azure VM. It's in Singapore, and I re-deployed it last night, so the numbers are still coming in, but it has a pretty tight distribution:
From a quick look, my own (stratum 2) server in the pool currently has an offset of just under 1/20th of one millisecond.
Regardless, thanks for contributing to the pool!
The bottleneck I kept hitting was the 65535 NAT translation limit on my Cisco router, at which point, load was quite manageable on the Pi.
It's extraordinary how much traffic one cheap device could service.
EDIT: AHA! Part 2: https://blog.pivotal.io/labs/labs/ntp-server-costing-500year...
My bad — I never wrapped it up. Thanks to the HN interest, I'll try to write Part 3 over the winter break.
The short version is this: it's gonna cost a couple of hundred dollars to run a 1Gbe NTP server in pool.ntp.org, but you can tweak the ntp.conf to save ~$100.
Incidentally, how is AWS dealing with the leap second next week? Google is going to have their time servers start to run fast around 20 minutes in advance of the leap second, so they're back in sync at 00:00:60 UTC.
If memory serves, Google's "smoothing" the second out over (I think) a 24-hour period. I don't recall the exact time period off the top of my head but it's much, much longer than 20 minutes.
EDIT: this post was indeed from 2014. My bad then. however the same issue started again two weeks ago (~17 dec 2016).
https://cloud.githubusercontent.com/assets/1020675/21468123/...
Note that inbound traffic which was steady at ~4k packets/sec spikes as high as five times as much. Also note that the snapchat traffic followed a circadian rhythm (much higher traffic during the daytime).