Semantic leakage could be just weakness of the model, and related to claims that they don't _really_ reason. Maybe more training could help.
Or maybe it's a more fundamental weakness of the attention mechanism? (There are alternatives to that now.)
Or maybe it's a more fundamental weakness of the attention mechanism? (There are alternatives to that now.)
No comments yet.