The problem is bias. Study after study has shown that those "casual chat" interviews are worse than useless at measuring anything at all.
Kahneman's book, Noise, has entire chapters on this problem. The only solution that empirically seems to work are a) interview panels and b) pre-defined standard rubrics with clear evaluation criteria.
Defining those rubrics is hard and the results aren't perfect. But when done well, you can get up to about a 70% correlation with on-the-job performance. Nobody is known to have achieved better.