> However, I think it's a bit odd to treat this type of use case as some sort of AI breakthrough that wasn't possible or wasn't frequently done in the past.
Classic computer vision is an utter PITA - especially when dealing with multiple libraries because everyone insists on using a different bit/byte order, pixel alignment, row/col padding, "where is 0/0 coordinate located and in which directions do the axes grow" and whatnot.
The modern "AI" stuff in contrast can be done by a human in natural language, with no prior experience in coding required.