1. inconsistent binary formats, float precision edge conditions, and unknown Endianness. Thus, assuming the Marshalling of some document is reliable is risky/fragile, so pick some standard your partners also support... try XML/XSLT, BSON, JSON, SOAP, AMQP+why, or even EDIFACT.
2. "dump and load" is usually inefficient, with an exception when the entire dataset is going to change every time (NOAA weather maps etc.)
3. Anyone wise to the 42TiB bzip 18kiB file joke is acutely aware of what compressed files can do to server scripts.
4. Tuning firewall traffic-shaping for a web-server is different from a server designed to handle large files. Too tolerant rules causes persistent DDoS exposure issues, and too strict causes connections to become unreliable when busy/slow.
5. Anyone that has to deal with CSV files knows how many proprietary interpretations of a simple document format emerge.
Best of luck, and remember to have fun =)