Rather than as a general parsing problem, isn't it possible to reformulate this in a different way that is much easier to solve by using the fact that these are strings describing flights we want to get destination and sources from?
So what you could do is take a general database of all flights (which is structured data and easy to work with) and then the problem is to label the unstructured input data with the probability it relates to a particular source, destination pair.
For example, there are no flights from York (lovely though York is) to Venice, so that mistake isn't one such a model could even make. Whereas there are a lot of flights to and from New York (including to Venice), and since that number is large it's a much higher probability labeling than some others.