The Shannon Limit (2010)
news.mit.edu
news.mit.edu
Almost all data sent electronically is first compressed until it is at entropy and then put into error correcting and line codes to handle channel imperfections.
The key is that the dictionary changes per-source.
This adaptive encoding would not really work for direct human interpretation, because you'd have to maintain that dictionary somehow in your head, separately for each source.
You'll note that our selection of acronyms and terms of art ("ROC curve", "tach", "JS") has something of this flavor, though -- and the lexicon is adaptive within each universe of discourse.
The first 17 verses of genesis, post LZ77, from http://www.infinitepartitions.com/art001.html:
001:001 In the beginning God created<25, 5>heaven an<14, 6>earth. 0<63, 5>2 A<23, 12> was without form,<55, 5>void;<9, 5>darkness<40, 4> <0, 7>upo<132, 6>face of<11, 5>deep.<93, 9>Spirit<27, 4><158, 4>mov<156, 3><54, 4><67, 9><62, 16>w<191, 3>rs<167, 9>3<73, 5><59, 4>said, Let<38, 4>r<248, 4> light:<225, 8>re<197, 5><20, 5><63, 9>4<63, 11>w<96, 5><31, 5>,<10, 3>at <153, 3><50, 4>good<70, 6><40, 4>divid<323, 6><165, 9><52, 5> from<227, 6><269, 7><102, 9>5<102, 9>call<384, 7><52, 6>Day,<388, 9><326, 9><11, 3><41, 6><98, 9>N<183, 5><406, 10><443, 3><469, 4><57, 8>mor<15, 5>w<231, 4><308, 5>irst<80, 3>y<132, 9>6<299, 28>a<48, 4>mamen<246, 3><437, 6>midst<375, 7><134, 9><383, 6><177, 6>le<290, 5><272, 6><413, 11><264, 10><429, 15>7<129, 9>mad<166, 9><117, 6><82, 6><348, 11><76, 8>which<215, 5><600, 10>nder<62, 14><115, 16><54, 11> ab<599, 3><197, 13><54, 9><470, 6><487, 7>so<169, 9>8<432, 20><108, 10>H<827, 5><397, 25><103, 9><405, 17>seco<814, 5><406, 10>9<406, 22><199, 8><235, 10><944, 7><428, 3>ga<439, 5><540, 10>toge<18, 4><45, 3>to one pl<820, 3><422, 10><604, 5>ry l<16, 4>app<981, 3><250, 8><474, 11><258, 12>10<258, 20><67, 9>E<1046, 4>;<638, 9><145, 6><234, 4><138, 8><86, 9><952, 13><75, 8><1018, 4>eas<853, 10><894, 6><883, 14><138, 9>1<290, 23><1179, 6>b<119, 5><1173, 3><11, 3>grass,<302, 7>rb<132, 9>yield<38, 4>seed<879, 10>fru<111, 3>tree<33, 10><19, 6>af<174, 3> hi<1229, 10>kin<57, 3>whose<69, 5> is<809, 4>itself,<1260, 10><148, 5><599, 23>1<1367, 16>brou<1082, 5><189, 12><58, 4><189, 4><181, 14><136, 9><154, 9><146, 7><204, 8><198, 19><175, 13><138, 4>i<1369, 10><184, 8><78, 14><401, 39>3<1160, 42>thir<753, 13>14<1460, 33>s<1155, 8><882, 10><1159, 15><780, 7><749, 3><1150, 11><100, 3><1031, 10>n<72, 4>;<769, 12>m<95, 4><361, 3><68, 9>sign<367, 7><22, 3><293, 3>aso<16, 12><79, 3><13, 7>y<430, 3>s:<192, 8><1486, 6><85, 15><185, 31><177, 10><126, 9>giv<1541, 8><573, 38>6<1343, 15>wo<562, 3><2001, 3><122, 7>;<906, 6><2019, 5><142, 7><288, 4>rul<1277, 14>d<1650, 12>l<1646, 3><45, 20><319, 6>:<937, 4><1452, 9>st<261, 3><647, 10>l<154, 11><1498, 10>s<278, 8><264, 33><256, 11><2099, 18><264, 5>,
For example the ideogram for "grass/straw/manuscript" when combined with "berry" makes "strawberry".
And that is exactly how English works. You take a basket, and a ball, and make basketball. Or grape, and soda, and make grape soda.
I'm saying we could make our communication more efficient, which you have not refuted.
*r* **u *k**?
means. Longer messages can be easier to decode, for similar reasons stated above in this thread. A lot of information is contained in the language outside of the character set, which is necessary for decoding. It asks a lot of the decoder to know everything necessary, but allows for relatively robust communication. *r* *o* ***y
or a** y** *k**
and you can see these are indecipherable and this is not a good example of "error correctable language".Error correction in language typically refers to understanding through context (i.e. other sentences), which would not be lost in this sort of compression.
All you skim.
And you skip.
Ate yak skat.
Add yak skin.
If you're telling me I can randomly hide 2/3rds of characters in any English sentence and you can still derive the meaning, I'm going to need a citation beyond "information theory" in general.But this is tangential to the point. I'm saying the context is preserved with compression, hence the meaning is preserved. I'm saying we can compress the amount of letters used and still get the same meaning.
Also there is a difference between having information and having perfect information. English can’t guarantee perfect information of arbitrary statements with 2/3 loss.