I don't have time to write an essay about this now, and I'm not sure I'd be the right person to do so anyway, but here are some thoughts I believe to be relevant.
Real Information Theory, as opposed to Data Transmission Theory, has to take into account the receiver. The transmitter has to have a model of the receiver, and then devise the data that will transform the receiver according to the intent of the transmitter.
For example, if you're a mathematician and I talk about the topological space obtained by gluing the edge of a mobious strip to the edge of a disk being the same topological space obtained by identifying antipodal points of S^2, you'll know what I mean. If not, I have to explain further.
Similarly, a computer program can start with a top level description of your intent, and a competent programmer can then complete the steps.
But a computer is not a programmer. When you describe your algorithm in sufficient detail for a programmer, you haven't given enough information for someone who isn't a programmer. A non-programmer will look at what you've written and then say "How do I do that?" Thus you need further instructions.
So you might say "We sort this vector of numbers by picking one at random, scanning down the vector to separate into those that are smaller, equal and larger, then recurse." and a sufficiently skilled programmer who's never met the Quicksort can now implement it.
But there are details missing.
You programs should (perhaps) be arranged so that the main routine is sufficient for a very skilled programmer to complete. Each routine called then is sufficient for a slightly less skilled programmer to complete, and recurse.
The problem is that the terminating case is the computer, and not a programmer, or even a person. Thus the terminating case is a long way down.
Perhaps there is a language design that will bring the terminating case higher.