Sorting Two Tons of Lego, the Software Side
jacquesmattheij.com
jacquesmattheij.com
The key insight here, to me, is that deep learning saves a lot of time, as well as being more accurate. I hear very frequently people say "I'll just start with something simple - I'm not sure I even need deep learning"... then months later I see that they've built a complex and fragile feature engineering system and are having to maintain thousands of lines of code.
Every time I've heard from someone who has switched from a manual feature engineering approach to deep learning I've heard the same results as Jacques found in his lego sorter: dramatic improvements in accuracy, generally within a few days of work (sometimes even a few hours of work), with far less code to write and maintain. (This is in a fairly biased sample, since I've spent a lot of time with people in medical imaging over the past few years - but I've seen this in time series analysis, NLP, and other areas too.)
I know it's been trendy to hate on deep learning over the last year or so on HN, and I understand the reaction - we all react negatively to heavily hyped tech. And there's been a strong reaction from those who are heavily invested in SVMs/kernel methods, bayesian methods, etc to claim that deep learning isn't theoretically well grounded (which is not really that true any more, but is also beside the point for those that just want to get the best results for their project.)
I'd urge people that haven't really tried to built something with deep learning to have a go, and get your own experience before you come to conclusions.
But seriously: I think you're overstating the case. I've enjoyed watching the fast.ai videos, and I think deep learning is a clear choice for some areas right now. Many others, no.
If you're working in a data-limited domain, in particular (i.e. most of them) there are probably better first choices. I know plenty of folks who have used deep learning to achieve no better than comparable results to the simpler methods they were using before.
But yes, the initial engineering costs of a deep learning system have come down a lot further than I thought before I watched your videos.
Most of the hate for deep learning on HN is usually against using it in titles as clickbait. (i.e. just simply taking a pretrained/predefined model and changing the source dataset without any optimization)
The in-depth explanations of mechanics discussed in deep learning as noted in fast.ai courses and this submission do justify the use of the deep learning moniker, though.
Deep learning is great if you're dealing with a hugely dimensional problem, and you have the data to train the model. If one of those things is not true (and usually, one of those things is not true), you're better off starting simple.
So even if I didn't have the data to train the model I had enough data to bootstrap the process and sometimes that's all you need.
Could you elaborate more on what has changed recently in the theoretical grounding of deep learning?
We've seen, in the last year or two, interesting results in nearly every area of theoretical research of deep learning, including generalization, optimization, generative modelling, bayesian models, and network architecture.
The truth is that an imagenet pre trained network will often work astonishingly well. The only other mark against deep learning is computational complexity, which I expect to go away soon (yours truly ;) )
It is possible that at some point to get to the last .1% or so I may have to embed features that were engineered but right now there is no indication that that will be the case.
It bothers me how often otherwise rational individuals continue to be susceptible to survivorship bias.
Of course you'd only hear about the successes, the failures are either embarrassed or know better than to tell you their story.
The people who tried it for several months, saw accuracy cap out at something like 60-70% for processes that need something like >90% confidence to justify expenses, and which proceed to get ignominiously shifted into a different team (or fired) for building what the higher-ups view as a massive waste of money at your salary and the other engineers see as a data-guzzling black box. No, these guys aren't likely to tell their story. Or, at least THIS story.
The story you instead get from these guys is how
>After having had some fun times doing deep learning and data analysis, I'm excited get into the big new field of $THING_THAT_WAS_HIRING
and all the other dross that modern vocal programmers use to mask any possible scent of failure and actual on-the-job difficulty experienced.
What I always try to do is to get a feel for the problem space trying different methods with as little investment as possible. That way you will - once you decide to go full power after a certain solution method - at least have a feeling that you are on the right path.
Dogmatically trying to shoehorn every problem into the toolset that you know how to use is a way to stay reasonably productive but it rarely leads to optimal outcomes, sometimes you simply have to learn how to use a new tool in order to get to the maximum.
I'm the last person to jump on new bandwagons, still have a dumb phone, don't use facebook and still run my own mailserver. Even so, when a tool has a significant and most importantly measurable advantage compared to the tools I'm already familiar with I'll adapt.
That's fine, but the thing is they once again plan to apply it on their badly maintained data warehouses (now often somewhere in their dusty hdfs stack) to build traditional predictive models. Customer churn, next best offer, propensity modeling, that sort of thing.
I try to keep telling them that with structured data sets, and simple classification problems, a random forest or God forbid, a logistic regression model, would do just as well given that you spend some solid time in feature cleaning and preprocessing. Am I wrong in this setting? Perhaps I should also jump on the deep learning bandwagon and start selling 10 layer binary Classification networks :).
Not disagreeing with you, but I do fear that many traditional industries will be jumping on board pretty soon without any of the actual use cases to warrant deep learning.
First people used to write classifiers by hand, but they found it's too tedious, unreliable and has to redone for each object you want to classify. Then they tried to learn to detect objects by using local feature detector and train a machine learning model to classify objects based on that. This worked much better, but still made some mistakes. Convolutional Neural Network were already used to classify small images of digits, but people were skeptical they would scale to larger images.
This was until in 2012 AlexNet came along. Since then performance of convolutional networks has improved each year. Now they can classify images with similar performance as humans.
A big reason is the tremendous increase in computing power available to the researcher for low cost. Most of these improvements depend on CPU-expensive training over lots of examples. In the past, the time to train a model or evaluate a situation would have been very high.
Another big reason is that datasets in a lot of these areas were fairly small, and the newer techniques tend to need a lot of data to train.
Another reason is that most previous researchers were focused on feature engineering, whereas modern techniques seem to move feature engineering into the ML system. This is a sort of conceptual change,
I don't really see it as "scandalous" in the sense that you'd expect people to have realized manual feature engineering wasn't the fastest way to get human-like results, or that computers were going to get faster for ML tasks, or that it would be possible to train deep networks, or that having good datasets against which everybody in the community can run and evaluate would be valuable.
It would have to be:
- less general
- a smaller process node
- possibly more than one chip on a board tightly coupled
- specialized data types
- very tight coupling between memory and computation
(so maybe memory on the chip)
- a slightly higher clock speed, say twice as high
GPUs are much too general but if all the factors above can be realized a factor of 10 in a PCIe add-in card should be possible.
Also, to get more training data, what about setting up a puffer to blow the part back on the belt and tumble it? If you could configure the loader belt to load parts slowly and stop after one is seen, you could automatically re-image the first part an arbitrary number of times by blowing it backwards before letting it move along and restarting to first belt to get another.
And question: do you normalize out color at any stage? As in, classify a black and white image, with a separate classifier for the color?
Because you really only need one camera (and two mirrors).
> Also, to get more training data, what about setting up a puffer to blow the part back on the belt and tumble it?
That's an interesting and novel idea. It probably will not work because the difference between the heaviest and the lightest parts are such that you'd blow most of the parts clear off the belt. Also the camera is sampling fast enough that the part would end up imaged in many positions without the ability to stitch the parts together again. But interesting.
> And question: do you normalize out color at any stage? As in, classify a black and white image, with a separate classifier for the color?
No, but I am considering using HSV or LAB as the colorspace to see if that improves accuracy or reduces training time to get to a given accuracy.
This is an approach that will work for parts that are so common that even after a few runs you have lots of them but the 'long tail' of lego parts is the vast majority of parts and they are quite rare.
Bravo!
[1]: https://www.amazon.com/XCSOURCE-Microscope-Endoscope-Magnifi...
> I simply don’t think it is responsible to splurge on a 4 GPU machine in order to make this project go faster.
2 things: 1. You can rent 8-GPU machines on AWS, Azure or GCE. 2. The incredibly wide applicability of machine learning means that an investment in hardware might not be wasted. Even if you only use the machine for this one project, if it helps you learn more about the field it will probably still be a good investment career wise.
Also, keep in mind that the dataset is still tiny and that a method that works for large numbers of images may very well fail if you only have a few tens to maybe 100 or so images per class.
Now, if you're changing the architecture (such as by adding additional categories of pieces), as I said, that's more tricky - what people usually do there is something like lop off the top layers and retrain them from scratch, possibly while freezing the rest of the NN (the assumption there being that the learned filters and lower layers ought to already be sufficient to classify a new category, which is reasonable since the lower layers tend to be learning things like lines and corners, all primitives which should be able to classify yet another square or rectangle etc).
Since this is the obvious response any reader familiar with deep learning would have while reading complaints about how slow your CNN is to train from scratch, it'd be good to discuss it in some detail what sort of finetuning you've tried and how it failed.
I was about ready to give up on it when I decided to try to bring up a net from scratch and that worked quite well.
Do I understand correctly that a checkpoint is just a snapshot of the model at a point in time? i.e. "Here are the probabilities of each outcome given the characteristics I have observed already."
Also, what does "fully converged" signify? Are there points in the course of training the model at which it is more appropriate to "save" progress than at other times?
In machine learning/deep learning, the decrease in training loss has major diminishing returns as training continues. Eventually, training the model hits a point where the loss barely improves each epoch/iteration. (fun visualization from one of my projects: http://minimaxir.com/img/char-embeddings/epoch-losses.png)
In some cases, the loss can stop improving entirely, or increase.
Pay-for-what-you-use large-scale learning will get models trained faster than local single GPU, then you run the trained model locally.
Can't wait for Skynet to go live! :-P
This is one part I didn't fully understand. In the previous post jacquesm said that the sorted Lego sets are more expensive than the unsorted one (and a fake piece destroys the price). So:
* Is he planning to make a few additional buck buying unsorting sets and selling them after sorting?
* Does he have a huge collection and is bored to try to find the pieces?
* Just a project for fun?
> After doing some minimal research I noticed that sets do roughly 40 euros / Kg and that bulk lego is about 10, rare parts and lego technic go for 100’s of euros per kg. So, there exists a cottage industry of people that buy lego in bulk, buy new sets and then part this all out or sort it (manually) into more desirable and thus more valuable groupings.
> I figured this would be a fun thing to get in on and to build an automated sorter.
He then impulsively bid on a ton of bulk lego on eBay and ended up with a garage completely full of the stuff.
Sounds like it started for fun, then spiraled out of control and is now a thing he would very much like to do to get all this lego the hell out of his life for at least enough of a profit to cover shipping it to and from his place, if not much more.
Would you be able to first make an inventory of all your available pieces. And then load a DB with (all?) complete sets and let the machine sort different sets in 1 bucket (starting with the most expensive set first?). Or how are you going to get your sets together?
One question: Wouldn't it have been easier to use a line scan camera and tether line aquisition to the belt's movement by attaching a rotary encoder which output would trigger individual line scans? That's the standard solution in the industry.
This is why you read HN. Interesting though had Jacques not made the original attempts I don't think the payoff above would have been as useful.
I'm curious about this wavy line--does it need to be specially encoded in any way or did you just squiggle the belt with a marker and let the software figure out how it lines up?
Easy cloud training: https://www.floydhub.com/
I look forward to seeing if you can push it further by leveraging faster hardware in the cloud.
I suspect the training time could cause you to lose interest in iterating improvements. But, how cool would it be to make the project even better =)
has anyone applied this sort of thing to voice recognition ? i see a lot of computer vision applications, but haven't found any audio classifiers amongst the CV articles
You should stick up a donate button, if you keep writing interesting articles about how it all works, I'd happily throw a few dollars towards the process.
I'd be interested to know what the HN sentiment is for this kind of behaviour; his only claim to these domains is that he got there first - he obviously has no intention of using them beyond selling to the highest bidder.
Recently I listed an unneeded but above average (trademarkable, keyword, .com) on HN as free to anyone who could use it, with the stipulation that they pass it on if subsequently not used.
I'd be glad to see more of that.