httpdiff – diff responses to two HTTP/HTTPS requests
github.com
github.com
diff <(curl -vs https://news.ycombinator.com/ 2>&1) <(curl -vs https://news.ycombinator.com/ 2>&1)
As a shell function: httpdiff () {
diff <(curl -vs "$1" 2>&1) <(curl -vs "$2" 2>&1)
}
httpdiff https://news.ycombinator.com/ https://news.ycombinator.com/ diff <(curl -Lvs https://www.google.com/ 2>&1) <(curl -Lvs http://www.google.com/ 2>&1)
On my machine that produces 33k of output. I like one liners and using the shell, but this tool was built not for fun because I was debugging things that were painful. $ diff -y -W $COLUMNS <(curl -Lvs https://www.google.com/ 2>&1) <(curl -Lvs http://www.google.com/ 2>&1)
Edit: furthermore, with curl you can use options like -b, -c, -d, which would all have to be reimplemented in your system."Those who don't understand Unix are condemned to reinvent it, poorly." – Henry Spencer
Using a 1970s TTY in the 21st century is dumb.
I would enhance the shell script a bit to output a few more commands, should I need them.
If it was a tool I used regularly I could sharpen it.
You can look at my shell script HTTP client if you like.
http://plan9.bell-labs.com/sources/contrib/maht/rc/httplib.r...
and some of my other plan9 code
The issue I've hit, that you may want to consider, is a suppression list of sorts. The ability to silence diffs on things that look like dates for example would be rather valuable.
(Also, it was PHP, which a lot of people hate.)
Years ago I found the Levenshtein distance is super helpful to determine how different the responses are, and used it as part of a black box web security scanner. You can do this just on the Raw HTML, but that's noisy and shows a number of differences. It's better to use an HTML-aware string distance function, that diffs just page content. I used that a channel for detecting blind SQL injection (in combination with some other things).
I also found that you can go a level higher, and use Levenshtein on just the HTML tag structure of different responses. By looking at page structure, and applying different weights based on the HTML tags that were added/removed you can group similar pages, which usually maps to the different functional areas/templates of a site. As in, you can say "these 5 pages are all product details pages", "these 10 pages are all blog posts", etc. Super helpful from a security scanner, since this could inform crawling/auditing choices and speed up audits. It also allowed us to say "you have a XSS vulnerability in your Blog comments form" instead of just saying "you have XSS vulnerabilities in these 100 pages".
Anyway, there is a lot of value in detected how different/similar various responses are. See some of Google's published work about detecting near duplicates for web crawling...
function httpdiff {
diff <(curl -L $1) <(curl -L $2)
} httpdiff https://www.google.com http://www.google.com/$ httpdiff https://www.google.com http://www.google.com
Doing GET: https://www.google.com http://www.google.com
Error doing GET https://www.google.com: Get https://www.google.com: x509: certificate signed by unknown authority
I'm sitting behind a pretty heavy proxy though; could be that. That or OS X(10.9.5) certificate store issue maybe?
I'm a bit of a newb though - how do I install this?
1. You need Go
2a. Type 'make'. It will build the binary and place it in bin/
2b. Use the go tool chain directly to build. Does the same thing as 2a.This tool looks great, but would not have worked with my particular use case, which was doing some migration of user data, and diffing the user accounts to make sure that they had changed in the expected way.
It might be an idea to add the ability to query the same host twice, but have a user input trigger when to test each host.
Why leave something dangerous lying around when /probably/ nothing is going to go wrong... until someone picks it up and decides to do something with it that was unexpected, when better alternatives abound?