The real problem in cases like this and other applications, as you and many others have mentioned, is that LLMs are basically correlation machines. They can find very complex, immensely-multivariate correlations in large data sets, and reproduce these correlations very well. But they cannot reason (so far) beyond these correlations in other to find deeper, less obvious causal relationships. They’re simply not trained to do that, yet. But it’ll come…!