Once you've determined what the information-content of a message is, then you can apply information theory to it. But different receivers can derive a different amount of information from the same message.
Consider, for example, that if somebody doesn't know English at all, then before receiving the message, their best guess at what it is is some probability distribution over all English characters (or sounds, depending on what we assume them to know), and after knowing the first part is "Help, I am" that distribution might not change much at all. Therefore, they derived very little information from this message.
Going in the opposite direction: keeping fixed the knowledge someone starts with, there is an upper limit to how sure they could be (even if they are logically omniscient) in completing the message (that is, a lower limit on the entropy of their probability distribution) - this is what you're talking about in your example. But this limit only becomes important under these constraints - for example, knowing more about the person who wrote the message can let you predict it better, and if predictor A isn't logically omniscient, predictor B can do better than it with the same prior knowledge, just by being smarter than A.