I really enjoyed reading this, got invested in learning about this idea, and then realized they they never actually show any evidence that this works? Like the text-based model they are experimenting with did a kinda counterintuitive thing that they thought was worth a section, and then they don't say whether or not it was right? The idea of aggregating chatbot context into a probabilities is super interesting and I really want to know if it's bullshit or not