EDIT: I think English would be least compressible. But I suspect that all written languages would compress quite a bit.
EDIT: I think English would be least compressible. But I suspect that all written languages would compress quite a bit.
The most efficient language is the least compressible language only in a narrow and arbitrary sense of efficient. There are many considerations such as what is efficient for the speaker, the hearer, redundancy to noise, efficiency with respect to particular purposes, etc. We can assume that natural languages will generally make a good trade-off across these factors, and searching for the most efficient language in one particular narrow sense is not very useful. Moreover, compression of text focuses only on surface form, completely ignoring the dimension of meaning.
My conjecture is that artificial languages will be more compressible because they haven't had time to get honed down, like English losing "thee" and "thou", that personal mode of address. Esperanto and Loglan are completely regular, which natural languages are not, and thus has a lot of use-cases where the regularity doesn't matter - they haven't had time to lose the mostly-unused features.
For better or for worse, compression of text only uses the surface form to compress, because that's the level that compression works on - letters or bytes or some other unit. You can't compress meaning. Meaning doesn't exist per se: colorless dreams sleep furiously, after all. That is, you can use perfectly sensible words and letters and even legitimate syntax, and still create strings devoid of meaning. A document consisting of perfectly spelled words, and legitimate syntax, yet without meaning like the colorless dreams sentence, will compress identically to ordinary text with the same orthographic and syntactical validity.
It's been generally done, just by determining how many syllables are necessary to express an idea; it also turns out to be pretty easy to measure IIRC, because people express ideas at a constant rate across cultures, so if you speak a less efficient language, you speak faster, and if you speak a more efficient language, you speak slower.
Romance languages aren't very efficient; English is pretty efficient, and again IIRC, the most efficient common language is Vietnamese, which can express ideas in about 2/3 of the verbiage that English needs.
I did find "Clustering by Compression", https://arxiv.org/abs/cs/0312044
The emphasis in this paper was not on "efficiency", but rather on phylogenetic grouping. The authors used the "Universal Declaration of Human Rights", which apparently has 52 translations as the subject text. See Fig 13 and section 5.2 for details. Instead of just compressing and normalizing, they developed a normalized compression distance, which involves concatenating two texts, compression, and division by the largest of the two texts, as compressed alone.
I think that using the same encoding of all the languages and some kind of normalizing would wash the results of differences due to encodings.