Lon.gs: A URL shortener in C
lon.gs
lon.gs
Cool hacks are cool I guess but from the sounds of it, the C code is a bit scary. If you had to build this, why not build it in Rust? At least it wouldn't be so terrifying from a security standpoint and you'd still get whatever performance is supposedly needed.
It really isn't a project that I'd recommend to anyone, unless I wanted a shell on their machine. I'm quite distraught that it is getting this many (seemingly blind) upvotes.
https://www.reddit.com/r/C_Programming/comments/4p5ung/longs...
(After a few minutes of poking)
A segfault only happens when you try to access a virtual memory address not mapped to your process, or try to write to a page mapped to you but not mapped as writable.
Remember, this is C. There's nothing to stop you from writing past the end of some memory like an array as long as whatever memory you're writing into still belongs to you. If you write far enough, you'll eventually walk out of your mapped memory pages and trigger a segfault, but if you don't stray toooo far, you can overwrite some important data and absolutely nothing will complain. When people say C isn't a safe language, they mean it. C will let you get away with murder.
And it turns out that there's some very important data you can overwrite this way.
> How would you get a bash shell?
Let's say there's a function foo() that writes some data to a buffer on foo's stack, and I can control what data it writes (because it's data from a text box on a webpage, let's say). And foo() is buggy and doesn't validate that in all cases the data I control fits in the bounds of the buffer it writes to.
I can then overflow the buffer with my data, and take advantage of that to overwrite the return address of foo(), because the return address for the function happens to exist past the end of all the stack local memory for the function. When foo() returns, it will jump to the address I wrote in there, instead of where it was supposed to go back to. And as long as that address is in the process's mapped memory pages, again, nothing will complain.
To get a shell, as part of the overflow I either insert the binary data corresponding to the x86 instructions for something like a call to execve("/bin/sh", ...) and then have foo()'s return jump to the beginning of my instructions, or I cause the return to jump to some other code or library that will do that for me that happens to already be in place (there are more sophisticated versions of this exploit called Return Oriented Programming).
If want to read up on this: https://www.win.tue.nl/~aeb/linux/hh/hh-10.html
These days it's not _that_ straightforward with thing like ASLR, NX and heap hardening. You need some sort of information leak for ASLR, then somehow start controlling R/EIP (which may prove difficult if there isn't interesting things on the heap nearby), write a ROP chain, and then pivot into something more useful (if you can't/don't want to ROP your way to a shell).
I can see at least one buffer overrun dependent on database contents, and I wouldn't be surprised if there's public-facing vulnerabilities in this thing, but I don't want to spend another 5 minutes looking.
If the database could theoretically hold it according to its schema, the program should be prepared for the largest row size that is still technically legal. (Where "prepared" may simply mean that it throws a "This should never happen" error and aborts the request.)
Here's a contrived example:
At the user creation page the malicious user wants to get admin access.
They have their username be: admin' --
They do this because they are betting on a user named 'admin'.
Their password is: alwaysSanitizeInputs
So, a user is created named "admin' --".
Now, they go in to update their password.
The DB nicely pulls out their username, "admin' --", and it says:
UPDATE users SET password = SHA1('newPassword') WHERE username = 'admin' -- AND password = SHA1('alwaysSanitizeInputs');
So what happens? Well, the user "admin" now has a different password, and the Malicious user knows what it is. That's why you still need to sanitize.You get around it by parameterizing your queries, and not blindly trusting data that comes from the database.
Ideally a corrupt record only breaks the requests that read that record with an HTTP 500 or something similar.
https://github.com/riolet/longs/blob/master/longs.c#L238 looks like an example of a buffer overflow.
Etc.
But that said, building a proper HTTP stack is not trivial.
If you want to use C language, then why not create Nginx module?
Nginx already solved the hard problems:
* HTTP parser
* Distribute work via event loop on multiple workers
* Useful load balancing strategies (not as great as HAProxy, but i am satisfied with it)
* Serious effort in dealing with CVE
* Widely used and battle tested
Here's a fine guide on how to write Nginx module in C: http://www.evanmiller.org/nginx-modules-guide.html
q3k@nihilism ~/Projects/longs $ python2 -c "print 'GET / HTTP/1.1.\r\n' + 'a'*2000 + '\r\n\r\n'" | nc 127.0.0.1 1337
q3k@nihilism ~/Projects/longs $ PORT=1337 ./longs
*** Error in `./longs': free(): invalid next size (normal): 0x000000000155a5f0 ***
...
Sounds like a fun heap exploitation challenge. Almost CTF-like.EDIT: It also throws a whole bunch of warning when compiling. [moved this here from the first line after child post mention]
The compilation warnings was an additional remark, I should've phrased that post better.
[1] - https://github.com/riolet/longs/blob/master/wafer.c#L334-337
I don't have much experience with C outside of embedded systems, so my security practices in C are probably less than ideal for PC/server based code.
$ curl -i http://lon.gs/abf
HTTP/1.1 301 Moved Permanently
Location: foo
X-Evil-Header: evilvalue
there were also examples of XSS and data URIs.(they claim to have fixed this elsewhere in the thread, but I guess some of the "evil" URLs still work)
So to sum it up - a great post, excellent work with a tool that has a lot of potential for specific scenarios. When deploying if using things like Cloudfare Edge and giving it a bit of Productizing (my day job is as a Product Manager for a global ERP business) then this could be a hit (pun intended :-J ).
It feels more like a warning to avoid WAFer, as opposed to a good advertisement.
And when your URL shortener goes down, all links are dead. Shortening URLs is the last thing you want to do when you need to preserve pages that might go down. Caching is a better solution.
That's not the primary use-case for it[2][3], but we're happy to support people shortening URLs if they want to. The reason there's room for this is that our URL shortening function has no ads, no third party tracking, no bloat ... no dark patterns.
That's worth something to some people.
[1] Oh By, Inc.
But it’s still a closed, proprietary database and there’s no way to decode an “Oh By Code” if the service goes down, which is the disadvantage number one you have with URL shorteners.
> The real utility are the easily recognizable codes, prefixed with "0x".
Why not using a prefix that helps differentiate your codes from hexadecimal numbers? How do you plan to get people to know your service if it’s not recognizable? Let’s say I put one of those codes on my business card, I still have to write somewhere that people should use 0x.co to access the content, which ruins the advantage of having just one code rather than a URL to my website (literally anyone knows how to open a URL in a browser).
Yes, that's correct. In this case, "Oh By" will continue running indefinitely. This is true for two reasons:
1) Building an extremely lightweight, ad free interface has a happy byproduct of ... being extremely lightweight. The infrastructure requirements for even a wildly popular "Oh By" are trivial.
2) Because Oh By is self-funded, there is no pressure from outside parties to break the pattern of an ad free, tracking free service with zero bloat.
If you're having trouble swallowing this, remember that rsync.net has been running continuously since 2001[1]. I think that's a credible track record[2].
[1] As a feature of JohnCompanies, the first VPS provider, and then as a standalone corporation in 2006.
[2] In fact, the rsync.net warrant canary turns 10 years old this year.
There's a 0% chance this will be running "indefinitely." You're just skirting the issue by making a promise you can't keep.
Well that's the real trick, isn't it ?
I'm happy to report that while Oh By itself does not display ads, we are a consumer of ads. We have a decent budget for sales and marketing.
I used to run a public one but it was getting abused for spam so I stopped.
http://www.archiveteam.org/index.php?title=URLTeam
(edit: and you can imagine how useful that's going to be once browsers start checking the Wayback whenever they see a 404... we're getting close to starting a trial of that with Firefox, with other browsers to follow.)
EDIT: Very strange, the redirect now goes to a URL I didn't enter... (http://www.sadfasdfasfdasdfsadfasd.com)
What are the api endpoints? docs? Can you view a list of all shortened urls? Can you delete shortened urls?
Can you change the base domain of lon.gs?
YOURLS was popular for a while, and I tried it, but I was concerned with running a not very popular PHP app even on shared hosting. At least Wordpress gets decent attention. I was worried of people compromising my own YOURLS instance against me.
http://www.cvedetails.com/vulnerability-list.php?vendor_id=1...
No idea about equivalents for nginx, or something that could run on Github Pages (these would probably be the same, given that nginx doesn't have an equivalent to .htaccess out of the box).
This is why you should never use C in networked code or when working with third-party data unless you absolutely bloody well have to: It's just too many ways to fuck up, and most programmers will.
If things like memory management are not the first priority when writing every line of code, then you shouldn't be using C, and that's really the only reason you should be using C, when there's a specific need for memory management.
Not surprisingly, this is exactly why the world has moved beyond C for the application layer, it's painful to have to think about that stuff constantly, so people just don't (this being an example).
Here's the short url of http://news.ycombinator.com --> lon.gs/amk
Going to lon.gs/amk redirects me to an overstock.com address for a specific product.
The site doesn't appear to have been hacked, at least there's no affiliate link in the URL I was sent to. It just appears the site is broken.