HTTP Client Performance – IO
blogs.atlassian.com
blogs.atlassian.com
The only time you'd use a NIO client, is on a server - for example, when building a caching proxy server.
The expected benefits are higher throughput of the proxy (lower CPU/RAM consumption allows multiple clients to operate in parallel).
Individual downloads being faster was never expected to be a benefit
This is true for the asynchronous network IO parts of NIO, but definitely not all of it. It also enables memory mapping and DMA, two things that can have significant impact on stand alone software.
DISCLAIMER: I'm not trying to imply that my way is "better" nor the only way HTTP GETs should be called. I'm only commenting on how I generally work and the trend I've anecdotally observed. Though I'd be interested to see if many others on here do prefer curl over wget for pulling binary files.
[13:17:18] i13:klausa:~% uname -a
Darwin i13.local 12.4.0 Darwin Kernel Version 12.4.0: Wed May 1 17:57:12 PDT 2013; root:xnu-2050.24.15~1/RELEASE_X86_64 x86_64
[13:17:19] i13:klausa:~% curl
curl: try 'curl --help' or 'curl --manual' for more information
[13:17:22] i13:klausa:~% wget
wget: Command not found. which {curl,wget}
to demonstrate that curl is in $PATH where as the shell cannot find wget (since we're talking about base installs it's safe to assume that the aforementioned utilities should be in $PATH by default - if included at all).This is how I ended up consistently prefering curl:
Last time I looked closely, I had better control with curl, and curl had better support for stuff. For example HTTP 1.1 was supported by curl, but not wget.
This makes curl a far better choice for testing HTTP servers.
And next, since I now know curl pretty well, and have it installed all over the place, it is my go-to HTTP command line client. I don't know of any reason to use both, or any reason to prefer wget, although in some situations maybe "wget <url>" acts slightly more like you want than "curl <url>". (Which is of course easy to work around by using "curl -OL <url>", if that's what you want, but convenient defaults are still convenient)
> "This makes curl a far better choice for testing HTTP servers."
Well yes. You're just reiterating what I said about curl being great for server testing. But this article is about download performance. When just downloading a largish binary file (eg a tarball), it's easier to just:
wget example.com/source.tar.gz
With that, you get a progress bar, it saves the file to disk and basically just does all the sane options you'd want for downloading by default (which I think is why most INSTALL / README docs tend to recommend wget in their instructions).Again, I'm not saying one is better than the other, but the convenience of wget's defaults (which I think you also highlighted later in you post) does tend to make me favour it for downloading binary content (with curl being my go-to for any server testing or text resources).
> "And next, since I now know curl pretty well, and have it installed all over the place, it is my go-to HTTP command line client."
wget also comes as part of the basic install with almost all Linux and other UNIX-like OSs as well. In fact wget actually pre-dates curl (albeit they're both >= 15 years old, so there's not really much between the two relatively speaking) so I'd have expected even your oldest live systems would still have wget.
I can totally relate to the "know[ing] curl pretty well" point though. There is definitely an argument for reusing the tools that you're already familiar with.
One popular UNIX-like OS that comes with curl and not wget is OS X. That might also drive adoption.
(I seem to remember having had to install both curl and wget on Ubuntu, but I am unable to verify this now)
I'm grateful for your input. Sorry if I came across as elitist - that wasn't my intention but in reflection, in does read a little that way.
> "One popular UNIX-like OS that comes with curl and not wget is OS X. That might also drive adoption."
Ahh interesting point. I guess that would have an impact.
And libcurl is used quite extensively to download all sorts of things, including binaries. Git, for example, makes heavy use of curl for http remotes.
Keeping wget off my servers has stopped countless attacks, yes the PHP scripts should be updated... but shared hosting isn't all that great as a sys-admin.
But most places where I need such a tool: My desktop (Mac), the load-balancers (OpenBSD), the database servers (FreeBSD), the hypervisors (SmartOS), the NFS servers (FreeBSD), basically all the infrastructurey stuff, curl is the default. I wouldn't swear to it, but I don't think wget is installed on most of those systems.
https://www.gnu.org/software/wget/manual/html_node/Recursive...
"The APIs of NIO were designed to provide access to the low-level I/O operations of modern operating systems. Although the APIs are themselves relatively high-level, the intent is to facilitate an implementation that can directly use the most efficient operations of the underlying platform."
I'd be curious to see how these each behave when needing to manage concurrent HTTP connections, fetching files from an ultra-fast server such as Netty, nginx, or Undertow.
Say 64, 128, 256, maybe even more simultaneously.