My hypothesis is that either Consciousness is a series of frames or we can emulate Consciousness as a series of frames and that you can run this type of recursive self iteration input to an llm with a buffer. The reason for the buffer is that the context window is limited so you would drop out earlier stuff and hope that all the important things would be kept in the subsequent frames.
A further experiment was going to add a set of tags that represented <input> and <vision> where input was the user input interpolated through a python template and vision was an image that was described by text and fed into it. So that the llm at each frame would have some kind of input.
I lost a little bit of interest in this but this has maybe resparked it a little bit?