I found two identical packs of Skittles among 468 packs
possiblywrong.wordpress.com
possiblywrong.wordpress.com
All your math can be correct but if your assumptions are incorrect, your results may not be accurate.
But to write his simulation, he just needs to fill virtual bags with random colors and count them, no math needed.
This is how I feel about engineering.
If you design something, prove it out with logic and math. It will work.
Or you will learn where you were wrong.
In other words, if you design something and prove it out with logic and math, sometimes it will work and sometimes it won't...
And for 'discovered' things, there is no reason to make a mistake.
However, most of the time you are working on something that is New.
So I was just thinking about that today. I just bought one of those ridiculous 'incline' rowing machines; as you 'row' it lifts your body off the ground to provide resistance (in addition to the normal rowing motions) - I'd bet money this design lasts rather longer than the type of cheap rowing machine that uses gas pistons to provide the resistance.
There's all sorts of ordinary, everyday stuff that must be engineered to not come apart while using the absolute minimum of material (or the absolute cheapest and lightest of material; I imagine shipping was a significant portion of the cost to me)
I mean, that's the thing about engineering stuff. If you are willing to spend money on lots of high quality raw materials, even someone like me without much engineering background at all can design something that lasts just by massively over-speccing everything. The engineering skill comes in when you have to build a rowing machine that can stand up to vigorous use by a fifteen stone american, but they only give you a buck fifty worth of aluminum to work with.
I mean, I guess what I'm saying is that I look around my world and I see a bunch of stuff that must have taken a huge amount of engineering skill to make (I mean, to make at the price point that I got it at... to be functional in spite of the absolute minimum of raw materials) - would all that stuff be 'new' to an engineer?
If you're not minimizing something, you're not engineering.
If enough people were prepared to pay double to buy something with a 25 year warranty, I'm sure companies would step up and make those products. People don't want to pay enough for it though.
Paying more now for a longer lasting product depends on how much you discount the future (ie effectively your internal interest rate), and on how likely you'll think that you'll even want to use that gadged in the future. Eg the NES was probably ridiculously over-engineered for most people, since realistically they'd only need to not fall apart until you buy a SNES or Playstation.
If I enjoyed selling things more, I probably would have gone for something more expensive anyhow, but I think the light weight is worth a lot in my situation.
Also, specific to Engineering, there are countless tools that help making sure your product will have a high probability of working as intended, anything from CAD tools, simulation tools (which it could be argued are a more automated way of applying logic and math), QA processes, etc.
With enough experience you might be able to build an informal, tacit model in your brain, though.
The last one there was more a comfort level thing...people like to not fall down to the next floor, but they also like to not feel AS IF they'd fall down to the next floor...to you overdesign to remove the vibration.
9 times out of 10, it was that NVH aspect that drove the design.
– Donald E. Knuth (1977)
For example, know that most HNers are fairly informed individuals in tech. Thus, I trust that I am going to find a good deal of content and insightful comments without reading the article - because I can typically glean the gist of the article and the tidbits that would have interested me anyway, from the HN community.
Ha. If that were true, why are there hourly front-page articles about "facebook is leaking your info!!!" and "google is evil!!", or my favorite "bitcoin {is,isn't} a great thing!!" ?
At the risk of angering folks and being downvoted/flagged, I think either there is a major disconnect between the types of people on HN that upvote articles and the people on HN who comment on articles, OR folks here are, on average, not as informed as you think they are.
Commenting without reading the article is fine. Speculating about what's in the article without having read it is a problem.
that does seem a bit disappointing. I think anyone simply quoting the article needs to explicitly show that that is the case because it was very misleading.
still this article was fascinating to me, from the idea that someone would go to the effort to the results of such effort. then top off my fascination with the idea of trying elsewhere in the country should there be more than one manufacturing point or trying to buy up a production lot as seeing the results of that as well.
Now how many packs of M&Ms? There are six colors there. If you go with the peanut version it probably is worse because they you have the variability of the peanuts which would make the chances of encountering packs with more variation in just the number of candies.
HN has memes. Turning unlikely things into SaaS, Rust/Crystal strike force and bikeshedding are all too real here.
And now apparently not reading source material has that potential as well. Which is kind of hilarious.
Kind of? And with slightly higher average standards of discourse? And I believe this is actually a compliment.
In my experience, HN comments under article are almost always more useful and more informative than the original article. The same is the case with various subreddits. When I read, say, /r/SpaceX, I also immediately jump into comments, as there is better quality info there.
This applies to mainstream news stories in particular. On HN, there's a good chance you'll find someone who was - or knows someone who was - involved in the topic first-hand, and who then proceeds to debunk various nonsense a typical news story contains. That's a huge value-added.
> I think anyone simply quoting the article needs to explicitly show that that is the case because it was very misleading.
Sure, I think making it clear what text is quoted (and from where) should be an obvious rule. And it doesn't have anything to do with whether or not others read the article; it simply saves brain cycles trying to understand the comment.
There are memes here too, just not in the form of image macros, so the faux-intelligentsia here pretends they're better than those boorish rubes that frequent reddit
Pursue your passions, whether it is climbing a huge mountain or finding identical packs of skittles. It is what makes the world great and interesting.
The only difference is, I normally spend my time automating the process with computer vision (a skittle sorter/counter wouldn't be that hard to build; opening the bag is harder than identifying colors). And then I never really finish the project.
I was thinking that time could be saved by not sorting the skittles so nicely. Just take a picture of the skittles from each bag spread out randomly. You should be able to figure out the color distribution just by counting pixels within each color range.
"That's why we built Loppy the Chopper over there."
I ended up learning there is a whole area of computer vision in industry- relatively "boring" stuff like just looking at a line of bottles going by. I went on a tour of the money making machines in DC and they stream sheets of bills by industrial vision cameras to detect whether the bills are within QC.
Still amazes me people get paid to do this. To me it's just like a big fun game.
One thing to remember in the panic over automation killing off jobs is that handling the normal case is maybe 20% of the work. Handling routine errors is 80% of the work. Handling the errors that crop up once in a thousand parts - or dozens of times in this small sample - is why running lights-out is so difficult, and why we'll always need human operators and maintenance staff who can diagnose and correct unexpected problems.
That said, I'm at least 97% confident that OP, in manually counting and marking down these 27,000 items, mis-counted or mis-marked at least once. So I'm not feeling too bad about my pessimism regarding the automated solution. Even if it can't identify if a given sample is "gross" or not, it will add 1 to a counter reliably!
Also, one test of 468 bags that shows one result does not prove anything about the average number required. You'd need to run this test many times for that, and for that you'd need automation!
The only requirement was a hopper for the teacher to dump the skittles into at the beginning. She would have a different amount for each group that was pre-counted and our numbers had to match hers at the end within a certain % of error.
It was a month+ long project and really fun to see each team's solution in how they moved/sorted/stored the final product. Some teams used long conveyor belts, others used short stacked belts, etc. It was all really interesting and I thoroughly enjoyed it!
Designing a hopper was a recently project for me and it was really fun. I learned a bunch about mass flow and realized that there are all sorts of interesting industrial problems involving hoppers.
And even this histogram assumes a distribution of total number of Skittles per pack (that varies) that I had to guess at beforehand. In hindsight, the final sample distribution suggests that I probably initially overestimated the true variance, and thus also overestimated the expected number of packs I would need to inspect. In other words, this experiment arguably took longer than "average."
So you're right-- this experiment could have extended into 700 packs, 800 packs... and still have been consistent with the assumed model, but I would have simply been in an unfortunate 90-th percentile possible universe where it took much longer than "average."
Honestly thought I was going to read a "just cause" blog on machine vision and process automation, e.g. 3 months to develop a functional prototype and train the system, 3 days to process 468 packs...and automated repeatability at the end of it all.
I can't imagine doing this manually - though it'd take me more time to write a Machine Vision solution rather than just doing it manually. :P
1. load packet contents on transparent substrate 2. short vibration to randomly disperse 3. image capture top and bottom 4. ML processing 5. dispose of load 6. goto step (1)
Cameras for this class of application are cheap, but if 1 camera is an explicit design requirement, then a simple motor with 90-degree camera mount to capture top, rotate 180-deg, capture bottom is simple enough as well (albeit less reliable given the introduction of a precision electromotive element to the system).
If alignment is critical, then error budget can be appreciably extended by marking corners of transparent rectangular platter with distinct color/shape on top and bottom to serve as fiducials in computer-aided skew correction/alignment, compensating for inevitable drift over time; it could also supplement human validation of archived images so orientation is easily determined.
So you're saying a machine and multiple cameras?
Wouldn't really consider a discrete vibrator afixed to a solid platter a "machine" per se, but if you want to call it that, sure. I gave the problem 10 sec of back-of-the-envelope thought. 10 sec more suggests you can skip vibrator integration by selecting a more rigid transparent substrate, e.g. glass...Skittles would naturally rattle on that as a package is loaded. 10 sec more suggests you can constrain camera movement even further by using carefully oriented mirrors. How far down the rabbit hole would you like to go?
Why? There is zero need for this. I know it's one of the buzzwords of the recent years, but identifying colors? Come on.
Keeping with the context of the GP's remark, a solution is more about achieving process speed while minimizing error and human intervention in the face of uncertainty--i.e. single-shot at the package level, which is at least a 50x speed increase over individual mechanical sorting.
I take it you don't engage in strenuous physical activities far from grocers. For instance.
Btw, thats my next big overhaul. How would you want the data? Javascript HTML table?
edit: would love an explanation for the downvotes
Hence, building.
* empty a pack of skittles onto white paper and take a photo
* use some existing image recognition libraries to count each color
* export to csv or smn and do excel magic
It feels like an almost-trivial computer vision problem.
Counting the number of round colored objects on a white background is not a hard problem.
I think I've found a project for this evening...
If you're pretty good with Tensorflow, you could do it with an object detector. I've built absurdly accurate object detectors using the TF Object Detector tutorial and just a little hand labelled data plus some good synthetic augmentation.
For this project's scope and scale it's probably not worth automating unless you're going to be repeatedly running the process.
https://hackaday.com/2017/02/06/mms-and-skittles-sorting-mac...
To the author's credit though, doing it manually probably made more sense in his situation. It would take a few months to build the auto-sorter, which would've had to be tweaked to handle the misshapen blobs. The collection of Skittles photos is impressive too, and hopefully will end up in a Maths classroom.
It took about an hour of Googling OpenCV questions. Also, I've never actually used OpenCV before, and don't use Python regularly.
I haven't tested it against the full set of images. Even if it's not perfect you could at least use this to narrow down ones that are within some threshold, then manually verify they were tagged correctly.
I'll post the code shortly.
It's definitely not perfect but could probably be improved a bit.
There are a few where strawberry and grape or strawberry and lemon are swapped. I think better lighting would help.
> I think better lighting would help.
That also demonstrates the difference between having the possibility to optimize input until the output matches the known values and having to process "real life" images made under different conditions.
The later is certainly more interesting: e.g. I dream of having an OCR that is "always right" no matter how bad the image is.
What would it take to determine which brand (skittles, m&ms etc) and which sub-brand (tropical / peanut, etc) and how many calories are on the table?
In my head, you would determine the relative color difference of each piece to determine the brand, but I'm not 100% clear on how that kind of ai works.
With better lighting you could probably distinguish between colors of different types of candy, assuming they're different enough. If you can determine the type of candy and know how many calories are in each piece it would be easy to count the calories, of course.
FYI all of your new comments are being marked as [dead] for some reason.
Have you verified the author's claim that this is the only match in the whole data set?
My program isn't perfect, but I only spent a couple hours on it. I think it could be improved, and better lightning when taking the photos would help too.
However, I was curious how well it did so I compared it to the author's results. My program misclassifies approximately 80 out of 27,740 Skittles.
I searched for matches and it did indeed find the same match the author found:
(0, './skittles/images/334.jpg', './skittles/images/464.jpg')
(1, './skittles/images/103.jpg', './skittles/images/384.jpg')
(1, './skittles/images/111.jpg', './skittles/images/134.jpg')
(1, './skittles/images/118.jpg', './skittles/images/353.jpg')
(1, './skittles/images/139.jpg', './skittles/images/168.jpg')
(1, './skittles/images/152.jpg', './skittles/images/281.jpg')
(1, './skittles/images/158.jpg', './skittles/images/244.jpg')
(1, './skittles/images/198.jpg', './skittles/images/255.jpg')
(1, './skittles/images/198.jpg', './skittles/images/334.jpg')
(1, './skittles/images/198.jpg', './skittles/images/464.jpg')
(1, './skittles/images/201.jpg', './skittles/images/338.jpg')
...etc
(first row is identical, subsequent rows are one different)Since there are some errors, I did get lucky. However, if it's good enough, you could also manually verify ones that are close.
Anyway, personally I'd rather write code than manually count 27,740 Skittles, but I'm grateful the author did this because it's an interesting dataset to work with.
With hindsight (although up-front simulation bore this out as well), we see that we would have to go back and manually verify anywhere from 8% (off-by-one) to 25% (off-by-two); that's every fourth pack. Now consider how much more unpleasant that manual counting is, since the Skittles are haphazardly un-arranged in the image. In short, I'm unconvinced that an automated-- while similarly accurate-- accounting would be that much more efficient.
I'm not sure I understand why sorting is necessary.
> Now consider how much more unpleasant that manual counting is, since the Skittles are haphazardly un-arranged in the image.
It's actually very easy: my program outputs images annotated with each recognized Skittle circled with the guessed color: https://imgur.com/a/jlPWXRf You don't need to manually count, just make sure all the circles are the correct colors.
I'm also pretty confident you could improve the accuracy quite a bit by improving the lighting when taking the photos, and/or incorporating something more sophisticated like a neural network (probably trained with Skittles from several bags, separated by color+rejects, with many photos in different random arrangements)
I guess this takes me a while, and seems significantly more error-prone to me than the corresponding re-count when they are sorted. Granted, it's pretty easy when the Skittles are sorted as they are in these images. But how quickly can you scan this image: https://imgur.com/a/KpddGdH and check whether there are any errors? And how confident are you in that visual check?
You are right that this approach may be good enough to find a duplicate, which was the primary objective of the experiment. But I had hoped that this might also serve as a useful dataset for future student exercises, in probability, or even in just this sort of computer vision project... but I wanted to have accurate ground truth, so to speak. Inspecting your spreadsheet, it looks like this algorithm is still less than 95% accurate, even if we only evaluate the "clean" images with Uncounted=0.
It's a good book!
https://www.amazon.co.uk/Newtonian-Penguin-Science-29-Aug-19...
https://www.insidescience.org/news/physics-knowledge-can-til...
The teacher had us all "guess" what twenty coin flips would look like. The longest any streak any student wrote on their paper was maybe 4-ish in a row or something.
He then had us all actually flip the coins and record the results. One student had like 11 in a row, most hit a streak of somewhere between 5-8 of the same result.
Lesson learned, we're really bad at guessing 50/50 streaks.
(1/5)^59 = 1/1.73*10^41
Apparently ~200 million skittles are made daily, so at that rate, we might expect to get a monochrome bag of skittles after ~10^33 years.
Apple is culturally significant and Blackcurrants are banned in the US (due to then carrying a disease that would devastate one of the forests). So one is more important than lime and the other is totally unfamiliar.
https://possiblywrong.wordpress.com/2019/04/06/follow-up-i-f...
[1]: https://www.amazon.com/Skittles-Original-Candy-2-17-Pack/dp/...
Since this is a nerd site, the next step is to use this:
Build a Lego contraption to push some skittles through a sensor that counts them.
They do speculate that the number of candies per pack is not IID, i.e., that there are (anti)correlations from one pack to the next. But without knowing more about the packing process, and presumably also having some lot/serial number information for each pack, it would be pretty hard to establish this.
This is similar tot the birthday problem: How many people do you need to have a probability of > x% that two are born on the same day?
It's something like 50 people to have a probability of > 80%. You can conduct this experiment at a school, using each class as a sample experiment to see if there are two pupils born on the same day.
> As an aside, I think the fact that this particular concrete application happens to be recreational, or even downright frivolous, is beside the point. For one thing, recreational mathematics is fun. But perhaps more importantly, there are useful, non-recreational, “real-world” applications of the same underlying mathematics. Cryptography is one such example application; this experiment is really just a birthday attack in slightly more complicated form.
I would say one of the first great discoveries for a person is the exponential series (a real world examples: population growth). Another is the divergence of the harmonic series 1/n and convergence of 1/n^2 (my preferred real world example: pizza slices that converge to 1 pizza or diverge to infinitely many). E.g. give me 1/n slices for the rest of my life and I'll pay you $100 (-:
When travelling, I also have go-to experiments that I like doing (e.g., elementary proofs that the earth is round/spherical such as: great circles; N-E-S-W always at 90 degrees; shadow angles [Erastothenes]; seasons; etc.)
There are other things to investigate that are not really "proofs" or "combinatorial evidence", but equally interesting. One example is using music (esp. the piano) as a physical logarithm device. The music "sounds" additive but the frequencies are multiplicative.
I've just ordered a huge box of fun sized M&M packets and will try using computer vision to count them to copy this study from start to finish for learning purposes.
"Yeah, I learned from this experiment that I don’t actually like Skittles, which is probably good, so a lot of Skittles were bagged and handed off to relatives. "
A few hours with a blob detection tutorial would have saved hours of tedium.
There is no formal definition of "a potato chip bag". Hence it's possible to "just" create "a single potato chip bag"...
No reason, no monetary plan.
Having said that, this small sample is indeed reasonably consistent (or at least not inconsistent) with that iid assumption for the color of each individual Skittle. We would not expect to see any 80+% red packs even assuming that color was perfectly uniformly iid, because the probability of observing such a pack is so small (less than 10^(-19)).
However, still assuming this model, we should expect to see packs with very small proportion of reds... and we do, with one pack having just 3 red Skittles, for example. The entire distribution of proportion of red follows the assumed binomial distribution very closely.