Do you have a cheap permutation function for large n? It seems like you still have to do it in two passes if you do this.
One reason cited in TFA for the half-shuffled files approach is that it's easy to rotate old data out of and new data into the half-shuffled files.