Intelligibility was one of my biggest concerns while making the hack: singing can be harder to understand than speech. On the other hand, the vocabulary is limited. Phrases like "turn left" and "turn right" are easily distinguishable, and other unobtrusive audio cues could make it clear that the words came from the GPS. Over the course of the hack day, intelligibility improved dramatically as I tweaked parameters.
Mainly, this is an experiment, and there are lots of ways to improve on the first day's results to make it more practical. Instead of following the melody exactly, the pitches could be chosen from a limited range to be consonant with the song. That would split some of the difference between traditional GPS speech and these results. Also, Yamaha has apparently withheld their higher-quality Vocaloid voices from the free Canoris API, so there's potential for improvement there as well. (The documentation warns that the initial release of the free voices works better in Spanish.)
Even more practically, the GPS could use the timing information to slip in spoken directions at less distracting moments, like a human companion might.