HNHacker News
TopNewBestAskShowJobs

dpw

390 karma · joined December 20, 2008

http://david.wragg.org/
submissionscomments
dpw··on How we use HashiCorp Nomad
I'm an engineer at Cloudflare, and I work on Unimog (the system in question).

You are right that even balancing of utilization across servers with different hardware is not necessarily the optimal strategy. But keeping faster machines busy while slower machines are idle would not be better.

This is because the time to service a request is only partly determined by the time it takes while being processed on a CPU somewhere. It's also determined by the time that the request has to wait to get hold of a CPU (which can happen at many points in the processing of a request). As the utilization of a server gets higher, it becomes more likely that requests on that server will end up waiting in a queue at some point (queuing theory comes into play, so the effects are very non-linear).

Furthermore, most of the increase in server performance in the last 10 years has been due to adding more cores, and non-core improvements (e.g. cache sizes). Single thread performance has increased, but more modestly.

Putting those things together, if you have an old server that is almost idle, and a new server that is busy, then a connection to the old server will actually see better performance.

There are other factors to consider. The most important duty of Unimog is to ensure that when the demand on a data center approaches its capacity, no server becomes overloaded (i.e. its utilization goes above some threshold where response latency starts to degrade rapidly). Most of the time, our data centers have a good margin of spare capacity, and so it would be possible to avoid overloading servers without needing to balance the load evenly. But we still need to be confident that if there is a sudden burst of demand on one of our data centers, it will be balanced evenly. The easiest way to demonstrate that is to balance the load evenly long before it becomes strictly necessary. That way, if the ongoing evolution of our hardware and software stack introduces some new challenge to balancing the load evenly, it will be relatively easy to diagnose it and get it addressed.

So, even load balancing might not be the optimal strategy, but it is a good and simple one. It's the approach we use today, but we've discussed more sophisticated approaches, and at some point we might revisit this.

dpw··on Guédelon Castle
Wow, unusual topic for the front page of HN. I visited about 5 years ago. It was my wife's idea to go (it's quite far from other attractions, so you have to plan a visit), but we both really enjoyed it. The castle is the highlight of course, but there's quite a lot more to the site than that, and it's easy to spend a full day there.
dpw··on However improbable: The story of a processor bug
Thank you! We often use on Creative Commons-licensed images in our blog posts. We always include credit, but we owe a big debt of gratitude to the people who take these photos and make them available.
dpw··on Cloudflare discolors the Web
I spoke too soon! We're going to disable this change for a while. Sorry.
dpw··on Cloudflare discolors the Web
I work for Cloudflare.

Thanks for bringing this bug to our attention. We have just rolled out a fix. You might need to go into the CF dashboard and purge the cache for your site to see the fix take effect.

dpw··on Qanat
I saw lots of these in Morocco, between the mountains and the Sahara.

Well, what I saw were the regularly spaced mounds of earth at the top of the access shafts.

dpw··on Intel Is Preparing a Major Restructuring of Their Graphics Driver
"As of yet I don't have a clear picture what this new driver will look like once evolved besides hearing 'boxes, many fucking boxes mate', when being told about the increased abstractions of the multi-OS-focused driver design."

If only more discussions about software architecture were that honest.

dpw··on Docker Goes native on non-Linux OS with latest beta
So "native" means "not VirtualBox" now? Docker for Mac might be a significant step forward compared to the previous solutions, but it still involves a Linux VM. I guess you could say that xhyve is a native hypervisor, but that's a bit weaselly.
dpw··on Benchmarking Message Queue Latency
A few things about the article that made me think "hmmmm":

No mention of testing set-up. Was the test client running on a different machine from the server? What kind of machines? What kind of network?

Many of the charts have the same "ballooning" shape, despite measuring very different systems. I think this is due to the "attempt to correct coordinated omission by filling in additional samples". As I understand it, all charts but the first have this correction applies (and it does sound like it is applied by manipulating the data, not by altering the measuring method). To understand the effect this might have, imagine testing a system that has a single request queue by making requests on a regular schedule, say at 1ms intervals. And most of the time, these take much less than 1ms. But one request is an outlier and takes 100ms. What will the "corrected" results look like? The worst case will by 100ms. The second worst case will be 99ms. The third worst case will be 98ms, etc. On a linear horizontal scale, this would give us a linear slope at the right hand side of the chart. Change to a logarithmic horizontal scale, and you get a chart with the shape seen in many of the charts in this article. This makes it impossible to tell whether the worst cases are due to a small number of outliers or not. I believe that the correction is well-meaning, but I think the uncorrected results would be more informative.

The use of line charts is a bit odd. They are connected to the origin, which is obviously a fiction. They are also slightly smoothed - where steps are visible, the steps have a gradient rather than being a vertical line. Where the number of data points is low, this leads to odd effects: In the two 1MB charts, the right third of the chart is just showing the value of a single data point! A scatter plot might give the reader a more honest impression.

The logarithmic horizontal scale of those charts tends to focus attention on the worst cases. That's not unreasonable - in some contexts, that's what you really care about. But outliers might occur due to environmental effects like kernel scheduling, VM scheduling, dropped packets on a noisy network etc., unless you make an effort to prevent such things. And it makes it very hard to see the typical values on the charts for RabbitMQ and Kafka where the range of Y values is large. Can you tell what the median latency for RabbitMQ/Kafka for any message size is? It looks like about 0.5ms to me, but it's hard to read it from any of the charts.

The number of messages involved is different for different message sizes. You can see that from the way the 1MB charts are stepped, but the charts for smaller message sizes are smoothed. For 1MB messages, it look like there are 5k or 10k samples on the charts. For the smaller message sizes, probably far more. Were all the tests run for roughly the same amount of time? Tests run for longer might see more outliers due to the environment.

"The 1KB, 20,000 requests/sec run uses 25 concurrent connections". With the implication that other test runs had different levels of concurrency. So what were they? What was the impact of changing the concurrency levels while the message size/rate was constant?

Is it possible that the client program making the measurements was introducing any artefacts (for example, being written in Go, did it encounter any GC pauses?). It would be interesting to see the the results of measurements against a simple TCP echo server, as a control.

My criticisms may seem too harsh. It is too much to expect someone to expect weeks doing rigorous measurements, and the resulting article would be so long that hardly anyone would read all of it (sounds like academia!). Someone might say that I should do my own experiments if I think I can do them better; but I have a day job too. I don't want to discourage the author; I think it is good that the author did the work he did, and put it up for everyone to see. But when articles like this get linked on HN and read by lots of people, they can easily get regarded as conclusive. Ideas about the performance of various projects get established that might not be well-founded and can take years to dispel. So all I'm saying is, reader beware!

dpw··on Rocket Fiber Launches 100GB/s Internet Service in Downtown Detroit
The linked article states 100 Gb/s, i.e. gigabits per second. But the post title says "100GB/s" which suggests gigabytes per second. Worth correcting, because getting 100GB/s between two machines in the same rack would be quite an achievement today.
dpw··on Shaky: ASCII Diagram to PNG
Yeah. This bit me on the first ascii diagram of mine I tried it on. But other than that, very neat.
dpw··on Is the Network the Limit? Dealing with Weave, CoreOS and Azure
Indeed. Kernel-bypass networking is great when you are building something that just receives and send packets (and maybe does other things that can be moved to other threads within the same process using lightweight synchronization mechanisms). It's not great when you want to interact with other processes, because then you have to go through the kernel anyway. And if you have to go through the kernel anyway, the fastest approach is to do as much as you reasonably can in the kernel.
dpw··on Is the Network the Limit? Dealing with Weave, CoreOS and Azure
We did try TUN/TAP. It seemed to be slower than pcap.

Our answer to weave performance concerns is on its way: http://blog.weave.works/2015/06/12/weave-fast-datapath/

We're really interested to hear how weave fast datapath works out in all kinds of environments, which is why we put the preview out. It's a shame that Arjan's post did not include numbers for FDP, but maybe he'll be able to include them in an update.

dpw··on Ask HN: Self-doubt, inferiority complex is killing me. How do I fix that?
I agree with other posters who say that everyone makes syntax errors, and that relying on documentation is fine (consulting documentation, and not just banging out something that seems to work, is a mark of a good developer)

With that said:

Coding fluency is something that can be improved with practice. Google Code Jam has all the previous rounds available, and there are other similar sites. Do a round once a week, and hold yourself to the time limit. It will be unpleasant, and you might not think you do well, but no one will be watching. That context trains you to write correct code quickly, and you will improve with practice. You'll find that the coding fluency you gain helps your programming more generally.

dpw··on Ask HN: What I should do with my life?
Are you in London? If not, move to London - it's where most of the start-up jobs and interesting tech jobs in the UK are located.

Also in London, there are also lots of tech meetups. Find ones that interest you. Go to the pub afterwards, meet people with overlapping interests, and find out what they do and where they work. It might lead to finding a more interesting job, a mentor, or just helping to develop a sense of what you can achieve and how to go about it.

Being smart (in the classic CS sense) has little to do with long-term success in the world of work. Make an effort to develop your soft skills.

I think most people find it difficult to split their energy between a full time job and side projects (particularly if your also want a social life, relationships, to get exercise, have non-tech hobbies, and generally be a well-rounded human being). Form a habit where you work on side projects for a regular sustainable amount of time (e.g. 1 hour) every day. By making it a regular habit, it becomes easier to persist even when you don't much feel like it. And if you are not already an early riser, a good way to make this time is to start getting up an hour earlier than you currently do (and go to bed an hour earlier, naturally).

If your job was interesting, but it got boring, then maybe you are someone who needs fresh challenges to stay interested. If so, look for an environment that has that. And if you don't like doing stuff for clients, find a job that doesn't involve client projects.

dpw··on Pivotal Introduces World’s First Open Source, Enterprise-Class Big Data Suite
Has any Greenplum DBMS code actually been released yet? Please correct me if I am wrong, but as far as I can tell, there is only a commitment to release it at some indeterminate point in the future. Until that commitment is fulfilled, it is not open source.
dpw··on WeaveDNS – A distributed DNS service for a weave network
You make it sound like achieving what weave does would be straightforward with the existing in-kernel encapsulation mechanisms. That's not the case. Even if putting some parts in the kernel would make sense in the long term, getting it right in userspace first is not a ridiculous idea.
dpw··on AMQP and Node.js
amqp.node has a callback-based API if you don't want to use the promise-based API. But this guy reimplemented it, apparently because he doesn't like promises ("This library solves Issues 1 and 2 if you can put up with promises, but it still didn’t solve Issue 3").

He must really hate promises.

dpw··on What your framework never told you about SQL injection protection
The PHP world is so adorable. It's like a bunch of little puppies, running into each other and falling over and tumbling around.
dpw··on How old are you and what's your oldest code online?
I'm 38.

If you run Linux on an x86 desktop, you use code that I wrote in 1997: The MTRR support in http://cgit.freedesktop.org/xorg/xserver/tree/hw/xfree86/os-... (big chunks of that are unchanged in 16 years; I also collaborated on the kernel-side MTRR support, but that has been rewritten a couple of times by now).

It's probably possible to find online traces of code I wrote a couple of years before that. But I think that is the oldest code still in use.

dpw··on Who is Iran's Ali Khameini?
His name is misspelled in the thread title. It's "Khamenei" (and pronounced a bit liked "harmony").
dpw··on Decline in CO2 may be 'permanent'
The linked article is about a decline in the rate of CO2 emissions, not a decline in atmospheric CO2 as the HN title suggests.
dpw··on Red Pitaya: Open hardware instrumentation for everyone
I'm sure this is good for many things, but it won't make a great digital oscilloscope. A sample rate of 125Msps is not really high enough for a bandwidth of 50MHz in that context. The low-end Rigol oscilloscopes do 1Gsps.
dpw··on ActiveMQ: Not ready for prime time
And while we don't use github internally, people can and do submit issues and pull requests there:

https://github.com/rabbitmq

(N.B. we need contributors to sign a contributor agreement.)

David (rabbiteer)