ParentFull threadprobe·Do you think 3 is better than 1 & 2 as context gets larger? I think for smaller data sets its mostly fine no. It's an interesting bet by TM. End state does seem some form of continual learning (model weights update like dreaming)View on HN