The Unreasonable Effectiveness of Character-Level Language Models
nbviewer.ipython.org
nbviewer.ipython.org
I'd suspect some of the perceived quality at higher orders (particularly for the source-code example) is just coming from the transition graph becoming sparse and deterministically repeating long stretches of the input text verbatim.
EDIT: Oh, misunderstood. I know the RNN can do the parens-balancing, and that's why Goldberg said that parens-balancing was impressive, since with his method you'd need to add other hacks around it.