We changed RoPE's theta from 10k to 1m and fine-tuned with 16k tokens long sequences.
Curious, what led you to adjusting the parameters this way? Also, have you guys experimented with ALiBi[1] which claims better extrapolative results than rotary positional encoding?