1213486160 has a friend: 1195725856
rachelbythebay.com
rachelbythebay.com
I was debugging a library that was a 'native' library for a scripting language and the code seemed to have a much bigger running footprint than I expected. It kept allocating this odd sized buffer, a bit over 13,000 bytes in size. Walking it back to the scripting language interface to C code the buffer it wanted was '32' bytes long but the scripting language was passing it as a string so 0x3332 bytes long. oops! Reading hex and seeing ASCII is a very useful skill to develop.
Uppercase letters are 0x40 + position of the letter in the alphabet, so "E", being the 5th letter, is 45, "I", being the 9th letter, is 49, and so on.
Lowercase letters are 0x60 + position of the letter in the alphabet, so "e" is 65, "i" is 69, and so on.
That also means that you can swap case by flipping a single bit (XOR with 0x20).
Finally, digits are 0x30 + the digit's numeric value (including 0), so the digit "5" is 35.
(All of these properties were very intentional on the part of ASCII's creators.)
"Hello" is
01001000 8 (h)
01100101 5 (e)
01101100 12 (l)
01101100 12 (l)
01101111 15 (o)
And when you see all zeroes, it's probably 00100000, the space character.
01101000 is h
http://www.asciitohex.com/ try and play around with it, it's fun.
http://www.docklandsljc.co.uk/2016/06/unicode-cuddly-applica...
The specific slide regarding ASCII code points is here:
https://speakerdeck.com/alblue/a-brief-history-of-unicode?sl...
The standard US QWERTY keyboard does not quite follow this, though it is close (there's some insertions, substitutions ;)).
&[7] at 26, ([9] at 28, and )[0] at 29 are off by one in their current QWERTY keyboard positions. If we didn't have ^ and * where they are, then &, (, and ) would be in the right places to continue your pattern.
@, ^, and * don't fit the pattern at all.
I should also have mentioned this amazingly scholarly piece by Tom Jennings, that explains probably everything there is to know about where everything we've been talking about came from:
https://web.archive.org/web/20030201161943/http://www.wps.co...
Unfortunately it looks like someone else is now running wps.com so you can't get this directly at its original home anymore.
They used these four-char-codes for error codes, file formats, return values and other random things. After a while I could read and recognize the text from the HEX value of the integers. Even now I still see them pop up sometimes on iOS as error codes from random networking or filesystem errors.
or reverse engineering hardware or software :)
I found it to be also quite draining when I was reverse engineering an old proprietary software system to develop a tool to import/export data to it. Definitely put my brain in a weird place after spending hours pouring through network dumps and determining how it all worked. It did kind of feel like that scene from the matrix where what's-his-name is staring at the streaming code on the screen, though :P
Sometimes it can save a lot of time chasing down constant factors.
In octal a 16-bit word doesn't split up evenly into bytes, so even recognizing ASCII is difficult. For example, the string "AA" in hex is 0x4141, where each 41 is 'A' - pretty easy. But in octal, it's 040501; 0101 is 'A' but gets multiplied by 4 in the upper byte. (One thing in defense of octal: the 8080/Z-80 instruction set makes much more sense if you look at it in octal.)
https://news.ycombinator.com/item?id=13045558
In octal a 16-bit word doesn't split up evenly into bytes, so even recognizing ASCII is difficult.
That's true only if your hexdump is in 16-bit words; in bytes, it's just as straightforward: A-Z is 101 through 132, and a-z is 141 through 172. Incidentally, these are also where x86 puts the single-byte inc/dec/push/pop instructions.
$ echo 'ABC' | od -bc
0000000 101 102 103 012
A B C \n
0000004It is a good example of how language effects design, than the use of hexadecimal vs octal and its impact on computer architecture.
idle programmer mode enabled It seems that it would be possible to have special tools built and installed, which would claim the file types and have an icon with a 'no entry' symbol superimposed on the original type (ie Word icon with a red circle and line through it, for .doc etc..) .. the tool itself could be a simple program that just opened a notification saying '<tool> not installed'
If you see a file that's called "something.doc", can you tell whether it's a flyer for someone's holiday party versus another sample of that malicious RTF that's been going around this week? All the extension does is let Windows put an icon on it and dispatch it to the right application, and if the application supports multiple file formats it does the actual identification by the file header.
For most public programs, the trust boundary is generally assumed to be the process. Any I/O the process does is assumed to be untrusted; it could do anything. But anything inside the process is assumed to work as the language says it does, because the OS is assumed to provide memory protection that prevents other processes from tampering with it. (Some big companies go a step further and dictate that you're not to trust 3rd-party libraries unless the code has been specifically audited; this is generally a sensible practice security-wise, but a huge drag on developer velocity.) If you couldn't trust the basic machine operators, you'd never get anything done - you'd have to write sanity checks everywhere, and then you have no guarantee that the sanity checks themselves aren't backdoored.
For many internal apps, the network is inside the trust boundary. It's assumed that any network connection comes from a trusted source, because otherwise the firewall would've rejected this. And being able to assume this saves a lot in developer velocity; it becomes feasible to write one-off internal tools without the devs having to carefully audit all the I/O & cross-process code for vulnerabilities. If you didn't have this trust, most of these apps wouldn't get written, because the productivity benefit they provide isn't greater than the cost of writing a hardened, secure system. It's not just networking calls; if you can assume that your users are non-malicious employees, you also don't need to worry about XSS or XSRF, pathological regexps, DOS attacks based on large payloads, etc.
Title: Triple-Triple Redundant 777 Primary Flight Computer
[1] http://www.citemaster.net/get/db3a81c6-548e-11e5-9d2e-00163e...
(Unless of course some IT department has transparent proxies that try to be too smart)
And that's why some sanity checks are important (magic numbers on the protocol, size limits, etc)
This had some amusing side effects when we encountered some services we'd never seen before, like the port on HP printers that sends every byte straight to print... apparently expecting PCL or PostScript but if it didn't understand it, it just printed the ASCII. Came into the office one morning to find all printers out of paper and 500 sheets sitting in the output tray. Oops.
That seems like a design flaw. :)
Print a graphic was done by sending exactly every dot to it, I mean, 1 to put a dot on the paper, 0 to blank. Encoded as a byte...
People are supposed to do defense in depth, where there are multiple layers such that a compromise of a single one only leads to limited damage, but too often it's more like there's a hard outer shell surrounding a soft chewy center where they thoroughly and completely own you once the outer layer has been compromised.
Any messages that don't start with the magic constant just get ignored.
It's not really a security feature, and it doesn't have to be a secret value. It also doesn't mean you shouldn't sanity check the other data fields. The sole purpose is to quickly rule out data that's blatantly incorrect.
That's all that is needed to avoid people having to pull up a debugger to figure out why things are going bad in production.
Isn't that actually part the problem here? It's getting an erroneous size and trying to allocate a big buffer so it can read the data even though the data isn't really that big.
One solution might be a fixed header size, and a header checksum. Allocate space for the header, read what should be the header, including the checksum, and if the checksum is correct, then allocate the space requested for the data. A fixed size header doesn't really help unless you are actually checking that it's valid before proceeding.
Also, even if we do constrain this to memory leaks, I would say that leaks just make the problem worse, and a real bad bug that causes a persistent DOS rather than just an ephemeral DOS, but it's still a problem without them.
Actually allocating the memory prior to confirming you need it is bad because if you get enough requests that do that quick enough, you eat up all the memory. If you aren't freeing the memory, you just don't have to be as quick. Considering that that crazy value in question here that is being passed to malloc represents over 1GB of memory, it wouldn't take that many requests at all.
If you can't, sanity check against negative numbers, make sure you check the return value of malloc(), and set a timeout on actually reading that much data. If the malloc fails, close the connection. If the timeout fails, close the connection and free the resources. On the public Internet, you probably don't want to send an error message, since it's just exposing internal system information an attacker could exploit. (In debug mode, running internally, you probably do want to expose this.)
You'll be just as angry when you find out some application only lets you send it messages up to 1 MB. What a stupid restriction, you'll think!
If you expect that it might be an accidental ASCII character, the smallest one you are likely to accidentally receive is space. 32 x 16MB = 512MB.
If that's still not a large enough atomic message to satisfy, allocate 8 bytes for the size. If sneaky ASCII creeps in to the most significant byte, it's going to at least be 2^57, or 512 exabytes.
Plus, this is just the first layer of defense. It's an easy and cheap one that keeps you from overallocating, but it's not the end of the story. You still need to examine the rest of the message for validity, and throw it out as soon as possible if you find it's invalid.
And, finally, if you are taking the "allocate a buffer to hold the message" approach, and you might be receiving multiple messages at a time, you have to consider the possibility of receiving multiple spurious messages at once. The only thing worse than a process spuriously allocating a gigabyte and trying to parse stuff into it is a process spuriously trying to do that a few thousand times per second.
But you shouldn't be allocating 100MB+ upfront anyway. No matter how big of a message you allow.
If the HMAC/CRC doesn't check out, don't process the packet.
We ran into a similar issue when someone was trying to send a line protocol data to a pickle port.
Yes this format stinks.
On top of TCP often length prefixed protocols are used, which means some (e.g. 4) bytes of the length of a "packet" are sent first and the packet is sent afterwards. Thereby you can create the notion of messages and message boundaries on top of the stream oriented TCP.
In the receiver implementation you first have to read the length, then allocate a buffer for the packet and then can read the remaining message into the buffer. If someone will send a HTTP request to such a receiver implementation it will try to allocate the buffer with this number.
The most sensible way to avoid this is checking the first bytes against a maximum message size first. However I have to admit: In my very first implementation of such a protocol I also have not thought about this.
In fact, one could say that these are HTTP's magic numbers: 'HTTP/' for the response, and a few ('GET ', 'HEAD ', 'POST ', 'PUT ', and so on) for the request. IIRC, one trick web servers use to speed up parsing a request is to treat the first four bytes as an integer, and switch on its value to determine the HTTP method.