NTK-Aware Scaled RoPE allows models to have extended context without finetuning
old.reddit.com
old.reddit.com
NTK-Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation.
---------
This is a claimed improvement over the recent Superhot scaling blog post.