How We Built the ARKit Sudoku Solver
blog.prototypr.io
blog.prototypr.io
You can just find the squares, group the characters in the squares into visual equivalence classes, assign each class an arbitrary number, solve the puzzle in terms of those numbers, then fill in each empty square with the (average?) image of the equivalence class it matches.
This would allow you to solve a Sudoku puzzle with letters or WingDings instead of numbers, and the output font would naturally match that of the original puzzle.
But treating the symbols in the starter cells as arbitrary is ingenious, imo!
This is super cool, but I can't help but think that something is missing if it takes hundreds of thousands of examples of digits for a machine learning algorithm to be able to differentiate them. It wouldn't take a human child this many. The available machine learning algos are not using near the amount of information available.
> We use iOS11’s Vision Library to detect rectangles in the image.
Looking at https://github.com/gunapandianraj/iOS11-VisionFrameWork - this definitely doesn't touch CoreML
The nice thing is it detects "projected rectangular regions" so even if the puzzle isn't aligned with the camera it still works.
I do wish I had more control though; it runs into trouble sometimes and there's not much I can do other than apply heuristics afterwards to determine whether I should throw out the sample or continue.
Example of a bad read from Vision Rectangle Detection: https://imgur.com/a/RSpTG
Well, it's technically correct - it did find a rectangle :)
Interesting limitations to work around such as vertical planes vs horizontal, and focal length.
Not at all surprised they saw better performance with almost immediate payback by training models on their own $1,200 hardware than running in the cloud.
Very interesting they trained their own character recognition model and not only that but built their own custom crowd-sourced image labeling system complete with accuracy checks and review screens.
Overall, fantastic write-up!
I was surprised to read that IKEA had 70 employees working on their ARKit app! (https://twitter.com/DanielZarick/status/917472837295837185)
This whole thing (including the backend tools) took me about the equivalent of 1 month of full-time work (I was doing it mostly nights & weekends though since our games are what pay the bills).
I brought in one of my (excellent) designers from Hatchlings a couple days before launch to make the cool grid "scanning" animation and to do our branding and logo.
I don’t know whether the functionality is present (last time I checked, the app wasn’t available in ‘my’ App Store), but integrating the app with their inventory system(s) and translating it also can’t be free.
I had fun writing a Sudoku solver in Julia.
It's interesting to compare how far the interface/speed has come in 6 years if you watch the videos of how it was done then: http://googlemobile.blogspot.com/2011/01/google-goggles-gets...
(Goggles did not do the image handling on device)
"After the first pass I had enough verified data that I was able to add an automatic accuracy checker into both tools for future data runs (it would periodically show the user known images and check their work to determine how much to trust their answers going forward)."
The data I had available to mess with was the difference in width of the top of the puzzle and the bottom (with some trig you can determine its angle relative to the camera) and the projection matrix of the camera relative to the scene origin.
It's not perfect but it works better than having nothing at all.
If the aim was to simply learn new tech, though, then I get it. I am just wary of ML being a hammer used on anything even remotely resembling a nail.
Someone on /r/programming pointed me here: http://apollon.issp.u-tokyo.ac.jp/~watanabe/sample/sudoku/in...
But the app seems to already handle those Ok without doing anything special: https://www.dropbox.com/s/arfd03kr8ieczk5/recursive-solver-k...
I should note that’s 600k small squares so each full puzzle scan yields 81 small images.