Sure. I have a SaaS, EasyALPR.com that uses ALPR for parking enforcement that I've been working on for a few years.
I'm near a major release actually, replacing my previous products with something called Parking Enforcer, which I believe is the best mobile app / vehicle grouping tool in market. I focus on business parks with 300-2000 parking spaces to patrol. It has been in beta for about 9 months.
I have been working algorithms related to ALPR data set matching a lot, primarily in Python.
I'm not familiar with storm / spark. However, one issue is that that license plate reads are not 100% accurate. So you are looking for a fuzzy version of the plate.
Its possible for collections of plates to sort of fuzz-out as incorrect members become more distant from previous ones, causing new matches to join groups they should not. You can write stuff to handle this but the original point was ~"this problem is in analysis not ALPR" which I agree with.
As far as computation, a new plate may or may not have a group to join. If it has been seen before, you need to look for the group "most likely" to be the same car. This can mean iterating over a large data set to look for the highest probable group.
There are some tricks to cutting down on the data set under consideration. For example it is much easier to only look at possible sightings from this past week in this area than every sighting from everywhere. (which when assembling a national database may be necessary)
But in my experience, even in the tens of thousands of plates, doing grouping requires task queues and tricks for quick identification and notification of matches. A watchlist (blacklist) is might be short, but the grouping task over months of data can be long.
Some plates are VERY similar but not the same car! Sometimes the ALPR camera forwards photos of fences or vehicle grills that are not license plates at all, and those must be ignored but not at a threshold where good data is thrown out. When bad stuff makes it to the user, they have to be easily disposed of if they make it through.
Data cleaning things can require additional capabilities of analysis, like breaking apart improperly grouped vehicles, yet ensuring they do not rejoin "bad" groups. Lots of details to handle, and to an end user, it is painfully obvious when "matches" are not correct. They expect magic.
One other note is that I found building an efficient and useful model architecture for these purposes to be challenging. There are more details in how the raw data comes in from ALPR scans that have to be handled before you even reach an individual "Sighting."