Appealing results... do we think there is scope to close the accuracy gap against AR models? or even leverage the "Bidirectional Reasoning and Self-Correction" into an overall advantage?
There is no theory why the gap shouldn't be closed. At the same time, people are trying to make them work for years by now and it's never good enough. But they are competing with incredibly optimized architectures.