How I Made a Heap Overflow in Curl
daniel.haxx.se
daniel.haxx.se
"Reading the code now it is impossible not to see the bug. Yes, it truly aches having to accept the fact that I did this mistake without noticing and that the flaw then remained undiscovered in code for 1315 days. I apologize. I am but a human."
If Daniel happens to read this, thanks for your hard work, and I really don't think any apology is necessary. After all, the source was right there for any of us to read and review.
[0] https://www.buzzfeed.com/chrisstokelwalker/the-internet-is-b...
"This report seems entirely correct and it hurts in my soul."
and the blog:
"[...] shipping a heap overflow in code installed in over twenty billion instances is not an experience I would recommend."
Harsh to have to bear that responsibility for so little reward.
personally i've spent some time thinking about how in my life it is very easy for me to pay attention as hard as i can manage and still miss so many things. i feel like things like code review are a great way to work on this sort of thing (attention training, global and contextual awareness, etc.). for example, there is this place that i've been walking by regularly for a couple of years and someone maintains a nice flowerbed there and i pause to take a glance at it and have been doing so for about a year since i first noticed it. but a couple months ago i noticed that there is a companion flowerbed a few feet away that i just had never noticed because i was so focused on the first one. was that companion flowerbed there all along? i suspect that it was! but i have no idea because i didn't even know it existed.
i have tried to extrapolate this to knowledge across humanity in general. i suspect sometimes it takes just one person to notice something, to make an observation that informs their behavior, and it begins to disseminate to us all. sometimes it takes several tries and a long time to happen. and i wonder how much low-hanging fruit there is out there that just none of us have noticed yet. and i think about how integrating a simple, fundamental set of best practices already established by others who have been paying attention can improve the low-bar for us all.
i think that julia evans touched on this in https://jvns.ca/blog/2023/10/06/new-talk--making-hard-things... which was discussed here recently and so did dan luu in his https://danluu.com/p95-skill/ which has also been discussed here somewhere.
and, to conclude, i too have a lot of gratitude for stenberg and all the folks who've worked on curl, and many other parts of our internet infrastructure that i usually take for granted every day.
Sometimes the "Don't fix it if it ain't broken" principle is interpreted as "don't support it until breaks", and leads to nasty surprises.
While there are likely sustainability issues, I guess we will still continue to have these projects and hats off to everyone like Daniel working not just for themselves but also for their community.
---
Yes, this family of flaws would have been impossible if curl had been written in a memory-safe language instead of C [...]
The only approach in that direction I consider viable and sensible is to:
- allow, use and support more dependencies written in memory-safe languages and
- potentially and gradually replace parts of curl piecemeal, like with the introduction of hyper.
Such development is however currently happening in a near glacial speed and shows with painful clarity the challenges involved. curl will remain written in C for the foreseeable future.
Everyone not happy about this are of course welcome to roll up their sleeves and get working.
Including the latest two CVEs reported for curl 8.4.0, the accumulated total says that 41% of the security vulnerabilities ever found in curl would likely not have happened should we have used a memory-safe language. But also: the rust language was not even a possibility for practical use for this purpose during the time in which we introduced maybe the first 80% of the C related problems.
[...]
We repeatedly run several static code analyzers on the code and none of them have spotted any problems in this function.
I know about Gentoo and other Linux distributions that do build from source but it has been very uncommon for me to see a company building things from source instead of using apt, rpm, docker in this century. It's at least two orders of magnitude faster. I remember how in the 90s I had to download the source, scan the README for the dependencies, recursively so, then building the libraries, then eventually building the program I downloaded first.
Security wise, we were not reading the code, except for Nethack.
Even if the engineering teams do not build every dependency every time they build firmware (like busybox), they will at some point build all of the components of their system themselves (as supported by yocto).
Access points/routers/smart TVs/set top boxes are typically be based on ARM/Linux/Android, and should be supported.
And in any case, C++ is always a better option than C regarding safety, when nothing else is available.
Go's runtime can be a showstopper, especially for multi-language projects.
C++ may be better than C, but many people feel that Rust is even better.
Rust is already available in almost all scenarios, there's no need to wait for it.
While it's safe to assume that C gets a decent amount of use on every platform, you can't expect all platforms to be as well-supported as the major ones. Undoubtedly, some of those platforms would be listed as low-tier if the C compilers cared to maintain a platform tier list. But a platform being low-tier doesn't mean you shouldn't use that compiler there.
As for trusting Rust or C on niche platforms, C is so full of UB, platform-specific choices, and vendor extensions, that it's hard to ever fully know how well this or that project will work. Rust is much less surprising, if it works at all I expect it to fully work. I'd definitely pick Rust on niche platforms if I have the choice.
While not as safe as Rust, it is definitly much safer than plain C.
(edited to fix typo)
It's not just possible, it's been done. You can compile curl with rustls, you could for a time compile it with quiche, and work is ongoing to compile it with hyper. Curl is remarkably modular, none of those are mandatory.
> And Rust supports a tiny subset of possible C targets.
Gross overstatement. Rust supports the vast majority of devices that people buy today. Even if you ignore platform popularity and just count platforms supported by gcc vs rustc/llvm, there's only a handful missing from the later. And if you're talking about vendor-specific compilers, a lot of them don't support modern C or C++ either.
Some likely first choices would be hyper (http client/server), rustls (encryption), tokio (async scheduling). The Rust ecosystem is quite rich in protocols and codecs, it shouldn't be too hard to find most (all ?) of the crates you need, but there's still work needed to bring them together into one curl-like tool.
Note that Rust crates tend to be more focused than what you're used to in C, made to be composed together instead of used as a one-stop-lib. So your dependency tree would look much bigger than curl's.
Hyper is a pretty robust HTTP toolkit, and reqwest is a higher-level library on top of it.
"Most places want to build from source" is not something he considered; there is a brief "for users against the C implementation on the platforms you care about".
With that, curl can be used on the exotic architectures and the rust/other language version can be used on other platforms.
I know reversing the burden of implementation seems flippant, but it's pragmatic. At some stage, it's less community-wide work to support Rust on a new platform than to spend extra time maintaining/securing dozens of C codebases. Curl may not be making its Rust components mandatory anytime soon, but other projects like python crypto already have, to say nothing of projects written in Rust to begin with.
rustc_codegen_gcc is pretty close to ready, let's focus on getting it out of the door, and more target triples supported by rustc and llvm.
Also in embedded Linux I guess we are technically "building from source", when using a build system like Buildroot or Yocto. But Rust is already supported there and in fact already in use in our project (due to python3-cryptography).
That said, we can mitigiate it using mrustc to generate c. Does anyone keep a list of all the targets libcurl is getting built for?
* use a SOCKS5 proxy, AND
* you use the SOCKS5 proxy to resolve hostnames (which AFAICS is NOT the default), AND
* the size of the buffer was changed from the default (100kb) to something below 65541 bytes, AND (EDIT: Not correct: 'libcurl' re-uses the download buffer for this, which by default is 16kB, however it is said that 'curl' itself sets it manually to 100kb UNLESS you use --limit-rate)
* the SOCKS5 proxy is too slow to handle the request immediately (which however, as the CVE states, can usually be provoked if the attacker has control over the request rate). (EDIT: This is wrong, the CVE actually says "Typical server latency is likely "slow" enough to trigger this bug without an attacker needing to influence it by DoS or SOCKS server control.")
So to me, the attack vector seems very small. Am I missing something?
> the SOCKS5 proxy is too slow to handle the request immediately
the request seg faults immediately, no delays, no redirects
root@1aac5e228e16:/build/curl-7.74.0# curl -vvv -x socks5h://host.docker.internal:9050 $(python3 -c "print(('A'10000), end='')")
Trying 192.168.65.254:9050...
* SOCKS5: server resolving disabled for hostnames of length > 255 [actual len=10000]
* SOCKS5 connect to AAAAA...
* Send failure: Bad file descriptor
* Failed to send SOCKS5 connect request.
Segmentation fault
https://gist.github.com/xen0bit/0dccb11605abbeb6021963e2b1a8...
"The target buffer is the heap-based download buffer in libcurl that is reused for SOCKS negotiation before the transfer has started. The size of the buffer is 16kB by default, but can be set to different sizes by the application. The curl tool sets it to 102400 bytes by default - but it sets the buffer size to a smaller size if --limit-rate is set lower than 102400 bytes per second."
I thought with a buffer of 100kB, this bug wouldn't trigger?
I guess this is mostly relevant for software that runs on shared infra, sends requests to a url provided by an attacker (e.g. webhooks), and uses a SOCKS5 proxy?
Since DNS resolution is meant to occur remotely, this CVE would be directly applicable here.
if(!socks5_resolve_local && hostname_len > 255) {
socks5_resolve_local = TRUE;
}
This is really a bad idea. For people who use anti-censorship tool to protect privacy, this can leak their identity through DNS.Yes, the author shares your opinion.
It could have been impossible, but you never know if there is a way to escape language VM barriers. However, the author clearly ignored the DNS hostname limit stated in RFC1123 [1], which is hardcoded even in Java libraries.
> A host name in a URL has no real size limit, but libcurl’s URL parser refuses to accept names longer than 65535 bytes. DNS only accepts host names up 253 bytes. So, a legitimate name that is longer than 253 bytes is unusual. A real name that is longer than 1024 is virtually unheard of.
DNS is not the only mechanism for resolving host name to address, even if it's what's used 99.9% of the time today.
Host software MUST handle host names of up to 63 characters and
SHOULD handle host names of up to 255 characters.
I don't see where the RFC sets a upper limit on host name size. The DNS defines domain name syntax very generally -- a
string of labels each containing up to 63 8-bit octets,
separated by dots, and with a maximum total of 255
octets.
RFC1035 makes this more explicit [1].It can just as well mean, that the compiler will attempt to execute a proof, that memory is never accessed out of bounds, without well defined ownership and within the lifetime of the underlying object. Which is what Rust does, for example.
Of course it can also mean, that the compiler will then additionally add internal failure checks and safeguards at critical places (Rust does not do this, but it would be nice to have in systems where one might worry about in-register bit-flips (high radiation environments, like X-ray scanners), i.e. stuff not caught by – say – e.g. ECC memory).
When one of the best, friendliest, and most transparent C programmers of our time is writing these posts, we need to pay attention.
Given that said programmer wrote,
> Such development is however currently happening in a near glacial speed and shows with painful clarity the challenges involved. curl will remain written in C for the foreseeable future.
I'm curious what conclusion you'd like us to draw.