The curl-wget Venn diagram
daniel.haxx.se
daniel.haxx.se
Wget needed one option to enable resuming in all conditions, even after a crash: --continue
Wget's introdcution in the manual page also states: "Wget has been designed for robustness over slow or unstable network connections; if a download fails due to a network problem, it will keep retrying until the whole file has been retrieved."
I was sold. Even if I by some miracle managed to get all the options for curl to enable reliable performance over poor connection right, Wget seems to have those on by default and the sane defaults make me believe it will also have this expected correct behaviour enabled even for those error scenarios that I did not think to test myself. Or - if the HTTP protocol ever receives updates, newer versions of Wget will also support those by default, but curl will require new switches to enable the enhanced behaviour - something which I can not add after the product has been shipped.
To me it often seems like curl is a good and extremely versatile low-level tool, and the CLI reflects this. But for my everyday work, I prefer to use Wget as it seems to work much better out of the box. And the manual page is much faster to navigate - probalby in part due to just not supporting all the obscure protocols called out on this page.
Maybe I'm misunderstanding, but curl has exactly that feature, it's the `-C` flag. If you want retries, there's `--retry`. I find curls defaults pretty sane, personally, I wouldn't want either of those by default for a tool like curl.
Yes, I want retries. They should be the default for a user-facing tool. Try searching the curl's manual page for "retry". There are no less than 5 different interdependent flags for specifying retry behaviour: --retry-all-errors, --retry-connrefused, --retry-delay <seconds>, --retry-max-time <seconds> and --retry <num>
If I "just" want it to retry, surely there's a boolean flag like "--retry" that enables sane defaults? Nope! --retry takes a mandatory integer argument of maximum number of retries. Surely I can set it to zero for a sane default? Nope again: "Setting the number to 0 makes curl do no retries."
curl is a good tool if you know it through & through and want exact control over the transfer behaviour. I don't think it's a good tool if you just want to fetch a file and except your tool of choice to apply some sane behaviours for you to that end, that would probably make sense if you are a human rather than an application using a library.
But... it does, though. From the man page ( https://curl.se/docs/manpage.html#-C )
> Use "-C -" to tell curl to automatically find out where/how to resume the transfer. It then uses the given output/input files to figure that out.
I just tried it, works perfectly. I don't really see a difference between writing `--continue` in wget and `-C -` in curl. And the use case for specifying it is not so exotic, you might want to do range requests for all sorts of reasons.
Look, it's fine if you prefer wget's command line: I don't think retrying a request is a reasonable default for a tool like curl, but reasonable people can disagree on that for sure. But curl is perfectly capable of resuming downloads automatically, you're just (very arrogantly) wrong on that one.
> But curl is perfectly capable of resuming downloads automatically, you're just (very arrogantly) wrong on that one.
I've never claimed it doesn't. I've only demonstrated that the default options don't do it and enabling the behaviour is more difficult than it maybe should be for a simple tool. I fail to see the arrogance.
Yes you did:
> You must specify the offset from where it should continue
No, you mustn't, you can specify - and it does exactly what you want. The docs are very clear and even provide examples. At some point you should stop blaming curl for your inability to read a man page and admit that you were simply mistaken.
It is not an “obscure special value”. Not only is `-C -` (or `--continue-at -` for the long form) well documented in the correct place in the manual, `-` is a common value in command-line tools (e.g. when specifying that a tool’s input will be STDIN instead of a file).
IMO theres too much complexity for 'sane defaults' to not just be 'surprising behavior' for someone else's use case.
I am in agreement that wget has "sane" defaults i.e. it acts like a bot or web crawler, or basic browser. Curl has always been easier to get things done with, though. At least in the land of http requests.
With the modern web, sometimes it's easier to use a tool like Puppeteer from a custom script. Especially if the sites you're interacting with are using a lot of JS.
If a person thinks they benefit from 100 situps a day, I'm not going to disagree. And if they think there is some benefit in reading man pages, well, thanks to those who take the time to write all that ducumentation.
I've read through most of the man pages of every tool I use at least once, but it has taken me years, and I've done it incrementally.
(Seriously, there's a million nix flags, and only so many brain cells. ChatGPT's better than Google for simple "what's the magic incantation?" searches, and laziness is a virtue. If you don't want to be lazy that's fine, but I think you're missing out).
Almost all ISPs went through a scheme setup by British Telecom - you could either have free internet but you paid for your calls, or you could pay for your internet, and have access via a freefone number - so effectively flat-rate.
But the flat-rate option disconnected after two hours, on the dot. Which was hugely frustrating because we had a voicemail variant that was hosted by the telco, and let you know you had messages waiting by pulsing the dialtone. And my modem did not recognise the pulsed dialtone as a valid dialtone, and refused to connect until we called the number and marked them read.
Which lead to one of my most UK-centric retro stories. I tried to connect to the internet, and it refused to dial. I blew away my wvdial config, and it refused to dial. I blew away my ppp config, and it refused to dial. I grepped / for the error message and it didn't exist. I ended up blowing away my OS (and accidentally installing onto the wrong drive, and blowing away everything non-OS too), and it still wouldn't dial.
So I dragged the modem & extension cord to my mother's PC, and shot off a mail to my preferred mailing list (one hosted by John @ linuxemporium, my preferred source of mail-order distros), and swiftly received the response that in order to be certified by BT to operate on their network, one of the rules equipment had to obey was to refuse to redial the same number x many times. And that all I needed to do was power-cycle the modem. Which I'd done by dragging it upstairs to my mother's PC. And my own machine had been wiped twice over needlessly.
Aside, there was a lady named Helen on that mailing list who knew everything about everything, and is everything I aspire to be today. She had opinions on which harddrives best survived salt/sea air, and why they weren't deathstars. Just an incredible amount of lived experience. I miss mailing lists.
That was probably one of the biggest death-bringers for the BBS era, no more long distance calls to get to what you wanted.
> Which was hugely frustrating because we had a voicemail variant that was hosted by the telco, and let you know you had messages waiting by pulsing the dialtone. And my modem did not recognise the pulsed dialtone as a valid dialtone, and refused to connect until we called the number and marked them read.
Yeah, some VM providers in the USA did that too, and it similarly confused modems. It's called "stutter dialtone" here, and the usual fix was to put some delay elements in the dial string, which were commas for Hayes command set modems.
> and why they weren't deathstars
They sure did earn that name! I was so hesitant to switch to HGST for ZFS pools due to my 90s/2000s deathstar experiences. Wouldn't run them in production for a while, of course now that I'm over it and trust them as well as any other enterprise brand, they'll screw it up again!
curl cannot, AFAIK, do this. People usually suggest using xargs, which is a mediocre substitute because it waits for all the URLs to arrive before invoking curl, giving up any chance at parallelism between the command generating the URLs and the one downloading them.
ds@swann3:~# (for x in {1..100}; do sleep 0.1s; echo $x >&2; echo $x; done) | xargs -L5 echo
1
2
3
4
5
1 2 3 4 5
6
7
8
9
10
6 7 8 9 10
11
12
[... and so on ...]
If the xargs call uses -I then --max-lines=1 is implied anyway.If you replace echo with something that sleeps you'll see that the pipe doesn't stall waiting on xargs so the process producing the list can keep pushing new items to it as they are found:
ds@swann3:~# (for x in {1..100}; do sleep 0.1s; echo $x >&2; echo $x; done) | xargs --max-lines=5 ./echosleepecho
1
2
3
4
5
starting 1 2 3 4 5
6
7
8
9
10
11
12
13
14
done 1 2 3 4 5
starting 6 7 8 9 10
15
16
17
18
19
[... and so on until ...]
98
99
100
done 46 47 48 49 50
sleeping for 51 52 53 54 55
done 51 52 53 54 55
sleeping for 56 57 58 59 60
[... and so on until xarg's stdin is exhausted]
And you can stop the calls made by xargs being sequential too for more parallelism with the --max-procs option (or use parallel instead of xargs): ds@swann3:~# (for x in {1..100}; do sleep 0.1s; echo $x >&2; echo $x; done) | xargs --max-lines=3 --max-procs=10 ./echosleepecho
1
2
3
sleeping for 1 2 3
4
5
6
sleeping for 4 5 6
7
8
9
sleeping for 7 8 9
10
11
12
sleeping for 10 11 12
done 1 2 3
13
14
15
sleeping for 13 14 15
done 4 5 6
16
[... and so on ...]
(I adjusted max-lines in that last example because my current timings made things line up in a manner that made the effect less obvious, adjusting the timings would have been equally valid, in a less artificial example like calling curl to get many resources timings will of course be less regular, perhaps these examples can be improved by randomising the sleeps)I'm not sure what you would do about error handling in all this though, more experimentation necessary there before I'd ever do this in production!
You can specify multiple URLs on the same command in curl so using xargs in this way would do what you ask to an extent (the connection would be dropped and renegotiated at least between each batch) as long as you don't use any options that imply --max-lines=1.
With the --max-procs option you could be requesting multiple files at once which may improve performance over wget -i – though obviously take care doing this against a single site (with wget -i too for that matter) as this can be rather unfriendly (if requesting from multiple hosts this is moot, as is the multiple files-from-one-connection point).
For an artificial example, change
(for x in {1..100}; do sleep 0.1s; echo $x >&2; echo $x; done) | xargs -L5 echo
to (for x in {1..100}; do sleep 0.1s; echo $x >&2; echo $x; done) | sort | xargs -L5 echo
The sort command will, by necessity, absorb the list as it is produced and spit it all out at once at the end. xargs can still use multiple processes (if max-procs is used) to make use of concurrency to speed the work, but can't get started until the full list is produced and sorted.----
[1] An unnecessary sort in an ETL process causing the overall wall-clock time to increase significantly
echo '--url https://google.com/' | curl --config -For example:
$ curl -sSLOJ 'example.com/file name.txt'
curl: (3) URL using bad/illegal format or missing URL
$ curl -sSLOJ 'example.com/file%20name.txt'
$ ls
file%20name.txt
On the other hand, wget (without any additional flags) will produce a file called "file name.txt" for both URLs. Well, technically you also need to add a --content-on-error flag to wget because this example URL 404s.Just the fact that `wget url` downloads a URL and saves it makes it a winner for me in command-line use.
This is not about "sane defaults", but about use cases.
But the filename directive of the content-disposition response header is entirely coupled to the idea of a filename. Therefore, it ought to take precedence.
(Though both can be made to do the other thing in some capacity)
You do:
wget url://to/file.htm
and a file named "file.htm" appears in your cwd.Using curl, you would have to do
curl url://to/file.htm > file.htm
or some other, less ergonomical, incantation.For me, I usually want to download files and I'm usually not doing any more processing, so I tend to type in *wget" first.
wget is better for download lists, but I'd accept the argument that a simple shell script is similarly easy.
Yes. But the GP said by default.
> you would have to do `curl url://to/file.htm > file.htm` or some other, less ergonomical, incantation
… which begs the conclusion that OP is unaware of `-O`.
--remote-name-all This option changes the default action for all given URLs to be dealt with as if -O, were used for each one. So if you want to disable that for a specific URL after --remote-name-all has been used, you must use "-o -" or --no-remote-name.
alias curl='curl --remote-name-all'
wget "url://to/file.htm?uid=foo&q=bar&rnd=4"I'd prefer wget to be a bit more clever when handling URLs query strings though, but I guess changing this behavior now might break some scripts.
wget does have options to use the name proposed by the server, and so another option to remove the query arguments would be useful, and in line with those.
However, if the potential issues can be resolved with sane defaults, I think this would be a great new switch to add.
What?
Oh, I see now.
Do you work with many tools that can't work with files if they don't have the "right" extension? I thought that was mostly a Windows problem.
that "killer feature" for cat would be turn `cat file.html` into `cat file.html > file.html` which means if you actually wanted to cat instead of cp you'd also need `cat file.html -o -` kinda glad curl doesn't have that killer feature.
It's as if he treats curl as his mark on the world of IT.
> I work for wolfSSL doing commercial curl support. If you need help to fix curl problems, fix your app's use of libcurl, add features to curl, fix curl bugs, optimize your curl use or libcurl education for your developers... Then I'm your man. Contact us!
From Wikipedia’s wolfSSL page²:
> In February 2019, Daniel Stenberg, the creator of cURL, joined the wolfSSL project.
Given that, saying cURL is “a popular personal project that brings [Daniel] a lot of cash” seems like a bit of a stretch.
> HTTP PUT
wget --method=PUT --body-data=<STRING>
> proxies ... HTTPS
wget --use-proxy=on --https_proxy=https://example.com
Curl consistently has more options and flexibility, but there's several things on the right side of the venn diagram where wget does have some capability.
"""I have contributed code to wget. Several wget maintainers have contributed to curl. We are all friends."""
Curl is a general purpose request library with a cli frontend (also used embedded from other programs, or as a standard library API in PHP etc).
wget is my goto if I need to download a file now, with the minimum of fuss.
curl is used when I need to do something fancy with a url to make it work, or when I'm fiddling with params to make an API work/debug it.
1. wget can resolve onion links. curl can't(yet). You'll get a
curl: (6) Not resolving .onion address (RFC 7686)
2. curl has problems parsing unicode characters curl -s -A "Mozilla/5.0 (Windows NT 10.0; rv:102.0) Gecko/20100101 Firefox/102.0" https://old.reddit.com/r/GonewildAudible/comments/wznkop/f4m_mi_coño_esta_mojada_summer22tomboy/.json
will give you a {"message": "Bad Request", "error": 400}
wget on the other hand, automatically converts the ñ to UTF-8 hex - %C3%B1 - and resolves the link perfectly.I've searched the curl manpage and couldn't find a way to solve this. Please help.
I'm having to use `xh --curl` [1] to "fix" the links before I pass them to curl.
Don’t know what the advantages/disadvantages are, but it comes with the default install. It’s usually what I use.
This diagram is clearly and unapologetically biased towards curl. Feels strange that the author of curl doesn’t know what wget actually offers.
I regularly forget the order for the values for --resolve, try searching for that word and figuring it out quickly
I've been relegated to grepping a flippin' manual
https://daniel.haxx.se/blog/2023/08/26/cve-2020-19909-is-eve...
Previously I worked on an open source project that pulled in many third party libraries. Users would run their corpo vulnerability scanners on the project and find dependencies with open CVEs and demand fixes, not understanding that in our usage of the libraries, the vulnerability is not exposed.
I think in 4 years, we had users open roughly 50 issues like this, which corresponded to exactly 0 real world exploitable issues.
A central vuln DB makes sense for sysadmins, but too many make it the end-all-be-all.
Without Happy Eyeballs web browsers can be slow fetching some web pages, for some users, waiting for a request timeout on IP addresses that don't work before trying one that works, or working but with the slower IP.
It's called Happy Eyeballs because it improves the visible page load time in web browsers for many users.
Is there a feature matrix to Venn diagram converter?
(Deep down) on my To Do list is comparing Ansible, Puppet, Chef, Docker, etc.
Which ultimately means some kind of feature matrix, right?
With a converter, we'd get Venns for free.