so thats about %15 more compute per forward pass with 0 extra memory which is just nuts, so for a streaming or disk-based setup its just free better answers. def wasnt gonna think of this myself.
config layers overall delta math reasoning word problems
baseline 80 0.5391 +0.0000 0.5850 0.6357 0.3500
rys 87 0.5452 +0.0061 0.6706 0.6000 0.2723
cartographer_repeat_x2 92 0.7741 +0.2350 0.8455 0.8214 0.6000
looks like the model gets a second/third go at figuring out how to approach the problem and it gets better answers.i tried a matrix of other configurations and stuff gets totally weird. like playing em through backwards in that block doesnt make much of a difference / order doesnt seem to matter (?!). doubling each layer got a benefit, but if i doubled the layers and doubled that block there was interference. doubling the block where the model is architecting/crystallizing its plans improves reasoning but at the cost of other stuff. other mixes of blocks showed some improvements for certain kinds of prompts but didnt stand out as much.