Robdns – A fast DNS server based on C10M principles
github.com
github.com
fd = socket(AF_INET6, SOCK_DGRAM, 0); if (fd <= 0) {
This should be fd == -1. fd == 0 is valid.
https://github.com/robertdavidgraham/robdns/blob/master/src/...
https://github.com/robertdavidgraham/robdns/blob/master/src/...
check errors from sendto() & recvfrom(). you don't want to loop infinitely when fd is hosed.
https://github.com/robertdavidgraham/robdns/blob/master/src/...
Check return value of your mallocs.
https://github.com/robertdavidgraham/robdns/blob/master/src/...
Check return value of fd.
Checking the return value of malloc is a bad idea (just rig the program to explode when any malloc fails). But casting the return value of malloc is incorrect, and that code should be using calloc.
I am embarassed for you.
In his defense most programs don't test the return of malloc and don't set the abort on failure so most programmers would just assume that if your not checking your return values that you made a mistake.
I'll not argue the point further; the HN lynchmob is already beating at my door for daring to disagree with a site darling. Had I looked at your username before replying, I probably would have forborn from violating the bubble.
To find one, you're going to have to catalog every malloc() in the entire program and record some kind of recovery regime --- maybe degraded performance, maybe dropped requests, maybe fallback to some kind of pool --- for every allocation.
I am not sure why error-handling is a difficult concept for you.
Yes, it's suboptimal to have a server that receives a request with a 32 bit length field indicating 3 gigs of incoming data that the bombs out trying to malloc 32 gigs.
I'm saying: what's an example of a server that reliably, for every allocation of every piece of metadata, every strdup, every hash table entry, every connection object, &c &c, has a recovery regime so that it doesn't have to fail ever when malloc does?
One way you could find such a program would be to compile a candidate and preload a malloc that randomly (1 in every 100 calls per allocation size, for instance) returns NULL. See if the program (a) continues to run and (b) passes some simple unit test suite.
My contention is:
(a) Those programs will be hard to find, because most "production ready" Unix code does not have that property, and
(b) The criteria I'm talking about matters a lot, because (1) memory exhaustion strikes at totally arbitrary points in a program's execution, not just at the points where you're prepared to handle it, and (2) attackers can pinpoint exactly the allocation they want to have fail.
As an example of an approach that doesn't seem to work well: in "memqueue", the function that creates HTTP headers in responses allocates an array of iovecs. If that malloc fails, the header creation function returns -1. The function that calls the header-creation function, http_respond, catches that error and itself returns -1. Nothing ever checks the error return from http_respond.
(Also, the loop in which memqueue reads entire requests into memory by continuously realloc()'ing a receive buffer has an integer wrap bug in it, though it's probably not triggerable. But incorrect is incorrect, right?)
There are many daemons that require you to keep what you have in memory and just fail the particular action you are doing instead of throwing everything away. DB servers for example.
memory exhaustion is an error that should be handled like other errors.
in memqueue, http_respond logs the failed memory allocation(in http_cli_resp_hdr_create) and returns -1. there's nothing else I need to do in this case. The connection will get dropped without a response.
when reallocing, I don't see the integer wrap bug. Can you point me to the line?
What I am suggesting is to not overthink system error handling. Just handle it; aborting is one type of handling but not always what you want. Programs run in various environments and to guarantee a defined behavior we need to abide by the standard.
Also, your chunked encoding decoder seems to be using a signed strtol() routine to read an unsigned length variable. I could be misreading; I didn't look carefully.
memcached is a good example of a daemon that shouldn't abort imho, many other daemons could happily abort. https://github.com/memcached/memcached/blob/master/memcached... . In other parts of the code, aborting was a better decision. https://github.com/memcached/memcached/blob/master/memcached... . If you cannot preallocate conn structures why bother running?
It depends on the goal in the end. As long as it's an explicit decision and not relying on environment the behavior is expected to be defined.
For instance, I worked on an enterprise proxy where aborting on asserts wasn't acceptable. Why? Because the customer didn't want to interrupt his users even though in our opinion the proxy state was out of whack. This created a nightmare for us because it was hard to debug. We ended up fork()-ing and aborting on the side to debug the cores.
Malloc failures don't always have to result in aborting the program. Cases can vary.
My suggestion to you is abide by the man page and always check error conditions. Don't overthink the failure cases of stdlib & posix calls. You never know the OS/environment your program will run in, and the only common ground you have is the standard.
Malloc failures don't always result in aborting. The common alternative, especially in programs that have careful malloc return value checking regimes, is to occasionally cough up remote code execution.
Userland systems programmers should assume the conservative default of ending the program immediately when malloc fails.
A similar logic guided C++ into throwing bad_alloc instead of returning NULL on allocation failures. And a survey of modern C++ code will show you that most C++ programs simply allow themselves to terminate when bad_alloc happens.
Regarding mallocs, again you don't have to overthink error-handling. Just handle it and figure out how to deal with it if it happens.
C++ has nothrow and try/catch. If you want to catch a failed new or abort is something up to the user. One thing for sure is that C++ aborts on failed new instead of stumping on unallocated memory.
On the other hand, consistency is important. Some parts of the code handled fd correctly and malloc failures and other parts did not.
I recommend you applaud good programming practices and criticize bad ones. Defending/arguing bad programming practices(whether they are edge cases or not) doesn't help.
DNS uses UDP primarily. I suspect the author meant "UDP/TCP/IP" by "TCP/IP".
Years ago, I wrote a basic TCP stack for a honeypot research project. It is hard and incredibly complex. So this statement raises a number of concerns, and will need to be audited before being used in production.
Of course, then you find some weird version of Windows XP that this breaks :-(
I used to believe that. It is unfortunately not true. TCP packets will come in all shapes and forms, and all must be treated equally, which is what makes TCP stacks so incredibly complex to implement.
The Linux TCP stack is quite safe and fast. Especially with tight integration with NICs hardware (checksum offloading and the like). I'm a bit unsure what a custom stack in userland can provide that the standard kernel stacks don't have.
The upshot is this: going through the Linux stack, a DNS server is limited to around 2 million queries/second. Using a custom user-mode stack, it can achieve 10 million queries per second.
That may change, as some providers are starting to put smaller limits on their response sizes (to mitigate certain kinds of DDOS and response spoofing attacks). Of course permitting TCP 53 is required for DNS to work; as is permitting UDP fragments (which poorly configured firewalls often block too).
My long-winded write-up here: http://opine.me/cert-advisory-on-dns-amplification-offers-li...
The text below is from the RFC, and no it does not relates to zone transfers but to normal DNS queries :)
All general-purpose DNS implementations MUST support both UDP and TCP
transport.
o Authoritative server implementations MUST support TCP so that they
do not limit the size of responses to what fits in a single UDP
packet.
o Recursive server (or forwarder) implementations MUST support TCP
so that they do not prevent large responses from a TCP-capable
server from reaching its TCP-capable clients.
o Stub resolver implementations (e.g., an operating system's DNS
resolution library) MUST support TCP since to do otherwise would
limit their interoperability with their own clients and with
upstream servers.
Stub resolver implementations MAY omit support for TCP when
specifically designed for deployment in restricted environments where
truncation can never occur or where truncated DNS responses are
acceptable.
And for the most important part :P Regarding the choice of when to use UDP or TCP, Section 6.1.3.2 of
RFC 1123 also says:
... a DNS resolver or server that is sending a non-zone-transfer
query MUST send a UDP query first.
That requirement is hereby relaxed. A resolver SHOULD send a UDP
query first, but MAY elect to send a TCP query instead if it has good
reason to expect the response would be truncated if it were sent over
UDP (with or without EDNS0) or for other operational reasons, in
particular, if it already has an open TCP connection to the server.So I can't imagine why records would all of a sudden exceed 512 bytes on avg either.
How are there any fewer RTTs with TCP DNS than there would be with UDP? I'm not seeing the efficiency here.
RTT might not improve at all. But lag might. Scripts often make the mistake of asking for information when they need it. Instead of before they need it, so it will be ready when needed. The suggested approach would pre-load the DNS info and might reduce lag.
In fact my own UDP protocol does something similar to Nagle as well. There's no good reason UDP protocols can't pick and choose what features they include. But most don't.
Even the reliability argument doesn't make sense. Yes, TCP is "reliable". But so is UDP DNS, and in exactly the same way: if a request or response is dropped, it's retransmitted.
Nagle, for what it's worth, is an HN contributor. You could just ask him. :)
{edit} though performance isn't really about wire time or packet size - its about cpu time on either end plus buffering. Including router time since that's a cpu in the path.
There's also some work showing that you can achieve very high performance with kernel TCP: https://www.usenix.org/conference/osdi14/technical-sessions/...
A good general purpose stack is the 6windgate stack. I know nothing about it personally, but I know that a lot of people do use it successfully.
Yes and no. Most of the complication comes from extra functionality (segmentation offload, checksum offload, SACK) or from functionality which is required by the standard but not relevant for a DNS resolver (congestion control, window management, TCP timers).
If all you're doing is accepting a TCP connection, reading a small request, and writing a small response back, you can remove about 90% of the code from your TCP stack.
That aside: if I had to guess, this would be Robert Graham's 10th IP stack. He's been doing this (specifically) since the late 1990s.
1. Path reachability, MTU discovery and MSS interaction
When sending outbound packets, you have to correlate incoming ICMP error messages in case they signal a problem. If the problem is that the packet is too big, you have to figure out what the MTU really is (which can take repeated attempts), so that you know what MSS to use (for TCP, or fragmentation boundary for UDP). If the path is unreachable, you have to remember that too. In both cases, you need some kind of global book-keeping so that you can do the right thing across connections. Some protocols (like active FTP) implicitly rely on MTU discovery on one connection signaling the MSS for another connection, so everything has to be path based, rather than connection based. Messy.
2. State management for error correlation
O.k., so you've figured out how to fragment an outgoing datagram and know what boundary to use, but how do you handle incoming error messages related to the fragments? Even for UDP, or other "stateless" protocols you actually do have to keep state so that you can correlate those error messages to the packets you sent. When the error message comes back, it will have the IP ID of the fragment, but nothing else is guaranteed.
This goes for (1.) too, but ICMP error messages can also be recursive and nested, and for a correct implementation you need to consider how to handle ICMP error messages that were themselves triggered by ICMP error messages. Several userspace stacks get this wrong, and can't correctly handle MTU discovery for UDP, or double-error correlation.
3. Heuristical and inconsistent caps on state
Many TCP implementations support selective acknowledgements and duplicate ack signalling, but what are their tolerances, just how much data can be retransmitted or handled out of order before you have the resend the whole window? there's no way to know, and if you get it wrong you can end up stalling a TCP connection for a significant delay. Unfortunately there are no simple limits, and in some cases the volumes are related to bandwidth delay products, necessitating some kind of integral control loop.
The problem with all of these is that they only show up "sometimes" and with particular networks or TCP stacks. I've limited these to interoperability issues - but there are other tricky complexities. For example, when building a TCP stack, do you optimise for throughput and so batch reads/writes of many packets - or do you optimize for a correct RTT estimate, and do things more synchronously. It's not possible to have both (at least with today's NIC interfaces); sometimes RTT is critical (e.g. an NTP implementation, a real-time control system or just any system that needs to rapidly recover from packet loss) , sometimes throughput is more important. Definitely complex.
http://snellman.net/blog/archive/2014-11-11-tcp-is-harder-th...
TLDR: There are TCP implementations that can't handle SYN retransmission which you have to interoperate if your TCP stack is the product.