Ask HN: How do looped transformers work with respect to context windows?
Posting here as I haven't found the answer anywhere.
I know that GPT Astra uses looping instead of reasoning tokens. My specific question is this: reasoning tokens pollute the context and makes compaction kick in quickly. Does this also happen to looped transformers like Astra?