dude thats sick! i tried it out and it works. theres a couple layers in there that are part of the voidy block that doesnt do much for the selected answer, so i narrowed it down to L48-53 where this model is mapping out its reasoning strategy, and repeated that twice, i got a big improvement over the original config (i chose some questions from atropos and claude code made some up so idk not like a real dataset).
so thats about %15 more compute per forward pass with 0 extra memory which is just nuts, so for a streaming or disk-based setup its just free better answers. def wasnt gonna think of this myself.
config layers overall delta math reasoning word problems
baseline 80 0.5391 +0.0000 0.5850 0.6357 0.3500
rys 87 0.5452 +0.0061 0.6706 0.6000 0.2723
cartographer_repeat_x2 92 0.7741 +0.2350 0.8455 0.8214 0.6000
looks like the model gets a second/third go at figuring out how to approach the problem and it gets better answers.
i tried a matrix of other configurations and stuff gets totally weird. like playing em through backwards in that block doesnt make much of a difference / order doesnt seem to matter (?!). doubling each layer got a benefit, but if i doubled the layers and doubled that block there was interference. doubling the block where the model is architecting/crystallizing its plans improves reasoning but at the cost of other stuff. other mixes of blocks showed some improvements for certain kinds of prompts but didnt stand out as much.