Most definitely no offence meant, but if you are talking about gigabytes in the context of JSON, XML, or any other text based format you are doing something wrong. And yes, in this case I will stand by my "lack of experience" concerning your entire team, I am sorry.
However you havent really addressed your use case anyhow but you just threw keywords around - gigabyte, IO, compressed, etc. You might want to elaborate on where you have to use XML files of the size of 50 gigabytes.
It doesnt really matter what you are using, Java was just one example. If the XML parser you employ has similar issues you must not be surprised if the outcome is similar. And you seem to be coming back over and over to software support (dependencies). Yes, particularly Java was poor when it came to that but as I said quite some time ago, that is an issue with that software not the document format.
I will disregard the 10M+ tests, but could you publish somewhere the results of the 5MB files?
Again, JSON and XML are way too similar to be anywhere close to what you described and aforementioned benchmark highlighted that. Yes, its dataset is average but I am sure you'll be able to extrapolate that for larger sets.
Apart from the apparent improper use for data of that magnitude, I could only imagine you used an XML parser that simply was not fit for the task and if you do that you shouldnt be surprised that it does not work.