Earlier this year I published a paper [1] that was partly about recognizing place references from text. Various NLP libraries and gpt-3.5-turbo were used in the comparison. The comparison was not the focus of the paper and newer LLMs are probably better, but in the specific case, gpt had a lower precision score than most of the tested NLP libraries and was also a bit more difficult to handle when trying to force machine-readable output.