Why TinyML is still so hard to get excited about
staceyoniot.com
staceyoniot.com
Sorry. What’s the ML piece here? QR codes are from the 90s and don’t use any ML I’m aware of…
> However, at the conference Warden told me that, while he’d quickly discovered that the model worked, educating people about new gestures was tough. “No one knows that these gestures are available,” he said. This makes sense. If you remember back to the launch of the first iPhone and its touchscreen, the first ads and demonstrations focused on things like taps and pinch-to-zoom. Those weren’t intuitive; they were taught.
And if that were true, that’s solvable just like with the iPhone by having tutorials when you boot your TV. I think what’s actually the case that the CEO doesn’t want to admit is that he’s having trouble convincing TV makers this is a useful model when they’re all going into voice-operated UIs. That and the BOM cost makes it unappealing.
And when the model gets it wrong (which it invariably does) you’ve got a laptop that failed to be in sleep in a bag (maybe you fallback to more primitive models that are foolproof like increased temp + lid closed = in bag). But seriously. If you go to sleep on lid close, putting it in a bag doesn’t really change your thermal envelope (should have happened on lid closed). And booting before your lid opens seems silly when Apple shows that it can be done near instantaneously. In other words, this seems like a PM developed feature for promo instead of good engineering being done.
> The second use case also helps with thermal management. In that use case, the laptop detects when it is on a hard or soft surface. If it’s on a soft surface, like a bed or a person’s lap, it will try to run cooler so as to avoid overheating.
Ok. Maybe this is interesting. But do you actually need ML or is it enough to define a thermal budget and recognize a solid surface can probably dissipate heat more quickly and anything beyond that doesn’t buy you all that much.
But, seriously, this is just a "look at me!" scream to get customers or investors. People do that with every single thing (doesn't even need to be an actual thing), and it's no fault of the thing at all. It's not a matter of being there yet or anything, it's just dishonest people.
I have an old Chromebook that runs Linux. I did have the problem that it would sometimes wake up in my backpack and run very hot. Eventually I found out that the plastic lid was soft enough that it would bend and let the screen touch the touchpad which would wake up the Chromebook. It was possible to disable this behavior but it was not simple enought.
...is there also a neat way to remotely transfer the recipe's ingredients to an oven or load clothes into a washing machine?
One of my buddies once joked about the smart homes that they allow you to unlock your front door while being anywhere in the world — but sadly, there is almost never a useful reason to unlock your house's front door from 2000 km away.
Maybe it's not something that affects you, or 80% of people, 80% of the time... that doesn't make it useless. Try to step outside your own life even a few steps...
… except unironically
You mentioned rekeying the lock after the work, so I assumed that key copying wasn't the issue as you found a solution for that.
That being said, it’s possible the ML piece is about reading multiple QR codes at once in different angles. I could see that requiring some ML. But still. It feels like a solution in search of a problem since they seem to be taking the “throw spaghetti at the wall and see what sticks” approach to building products.
Someone did make such a washing machine, from 2012: https://www.appliancesonline.com.au/academy/appliance-news/s...
Note the caption: "The washing machine's brains do the thinking for us".
But I could be convinced that our clothes would wear more slowly if the machines had more data about what they were washing.
1) Wash in cold water 2) Use less detergent 3) Set your dryer on the "low heat" setting.
Yeah, no thanks.
You must have missed the memo, everything is AI now.
Your toaster? It toasts your bread with AI.
Your microwave? It heats your food with AI.
(Yes this is sarcasm. "AI" has become a meaningless buzzword thanks to marketing and the media. Same goes for "machine learning".)
Presumably they are using object detection to recognize where the QR codes are located in the image? Which they then feed to the standard QR decoding algorithms.
> That being said, it’s possible the ML piece is about reading multiple QR codes at once in different angles. I could see that requiring some ML. But still. It feels like a solution in search of a problem since they seem to be taking the “throw spaghetti at the wall and see what sticks” approach to building products.
What I’m saying is that washing clothes doesn’t seem to benefit from any QR codes. And the cooking example is even more confusing because presumably you’d have one for the recipe, not per ingredient
> What I’m saying is that washing clothes doesn’t seem to benefit from any QR codes. And the cooking example is even more confusing because presumably you’d have one for the recipe, not per ingredient
Well, we are in the age of IoT. Every appliance now wants to be connected to the Internet.
Clothing has care labels on it. Washing machine should read that.
IMO, what is needed is a larger focus on resource consumption in competitions, benchmarks and rankings. For example, I like that you can rank some lists on paperswithcode by number of parameters [1].
[1] https://paperswithcode.com/sota/image-classification-on-imag...
Beyond local control, I gave other examples of embedded AI in other comments — or more specifically, what you can do with an embedded LLM. Namely, being able to have a better human interface for complex settings, and being able to reprogram protocols (or anything that is “software-defined”) for future-proofing.
We already have SoC that is functionally not so much different from modern microcontrollers. Depending on economics of scale, it isn’t that big of a leap of imagination to see a microcontroller which includes 4-bit vector ops like a GPU.
Can you explain? I don’t see how sending data to the cloud is a huge burden compared to say an EV or your AC unit.
Even if you put together a private cloud in a data center, it's still going to use up a lot more electrical resources compared to say, an iphone-sized usage, much less in a low-powered, embedded application.
Also, there's a tendency for our civilization, when we make efficiency gains with breakthrough technologies, to then expand our usage. We don't do a great job of actually reducing overall energy expenditure.
How does such a "private cloud" differ from using the "public cloud"? The two seem identical to me.
It's not even a binary at many of the large providers: you can have dedicated servers ("my own server") in a public cloud at AWS, for example.
Personally, I keep it simple. If I can physically hit a box and no-one can get mad at me, then it's private and mine.
If you build the thing you are talking about, you will create the excitement currently missing.
For example, I made a photobooth simple device based on a Pi and a regular DSLR; the person takes a photo by pressing a button, then the Pi checks the camera and sends the latest image in a web gallery somewhere. The problem is the button: it needs a remote. If wired, it risks destroying the whole apparatus if someone pulls on the cord; if wireless, it risks being lost.
A gesture-based trigger would be super cool; but having it run on the Pi through the DSLR risks damaging the camera, so it should run on a different device, cheap and not too power hungry, such as an Arduino.
What about some kind of MIDI-controller based on an Arduino that could recognize gestures?
I also made a webapp to learn sight-reading (babeloop.com) but it requires users clicking or tapping on the screen, which is not natural.
It would be cool to be able to listen to users reading notes out loud and detect if they're right or wrong, on the fly, locally in a browser or on a phone. A light general speech recognition model such as VOSK is 40Mb and is able to recognize most phonemes; but for music sight reading, there are only 7 syllabes, so one should be able to make a much smaller model? I don't know how hard it would be though to train my own model...?
Arduinos can recognize gestures, faces, etc. right now with well-established tech, no servers needed. I've had a couple of robots running around my place that do this sort of thing for a number of years now.
Tiny SoCs will likely need accelerators to run more interesting models. For example, ARM is working on a ML coprocessor to pair with their Cortex-M chips.
It seems that transformers have succeeded because they can be parallelized, and solve many of the problems that RNNs have. This has lead to the advancements we've seen recently with GTP. Big companies are able to throw their data centers at the problem and train impressive and massive models.
Part of me hopes that a non-parallelizable AI architecture might be discovered which performs even better. Perhaps the problems with RNNs could be solved in a way that doesn't parallelize? We would be so fortunate if a desktop computer could run an AI that's half as good as Microsoft's or Amazon's best AI. I would love to see the advantage of the data center removed.
Philosophically, this does make some sense. The wisdom behind sayings such as "adding more people makes the project even later" exist because the greatest intellects we're aware of (ourselves) do not parallelize well.
TinyML does result in much cheaper products, though.
Last week, we were trying to fine tune a 175b parameter model to take a natural language prompt with hopes of directly-outputting correct, domain-specific SQL. Now that reality has passed, we are looking at different paths.
As of this week, we are trying to hit everything with the binary classification hammer. Turns out you can train a model to output 1 of 2 possible tokens with exponentially fewer parameters, training items, machine hours, etc. The statistics available in binary classification are also incredibly powerful and the results are trivial to reason with.
Even if you need thousands of binary classifiers, their scale and granularity makes this a non-event or potentially an advantage.
The real integration magic with AI/ML is starting to look like a weird form of set theory. At this level of complexity, detecting (and potentially confirming) the user's intention is way more important than trying to draw a direct map from input to destination.
Such things as air quality management are good use cases. You can't use your phone to do that.
Why not? I am strictly against IoT in my household so I may be way off base, but why can't your phone control your air purifier?
For example, I have a dishwasher with a bunch of settings, can sense load, etc. It’s got a touch interface that works with wet hands. Or I can tell it to start with the usual settings, or that a particular load is a bit different. Same with the laundry, the pressure cooker.
It is less mind bandwidth when you got kids.
What I don’t want, is for my appliances to do is to phone home to the makers.
LLMs (if you don’t somehow trigger its insanity) can be far more capable than Siri. How do you get that into something more energy efficient than a high end gaming rig?
Something more hidden is using LLMs to reprogram machine-to-machine protocols. That might extend the lifetime of machines that have to talk with other machines, but it breaks planned obsolescence.
There are plenty of exciting product ideas. Whether they are exciting revenue generators are another thing entirely.
In the context of the article, an LLM is kind of the opposite of "TinyML" and not something most IoT devices could even handle.
Article aside, reducing energy use for models is one of the research areas for TinyML.
Thinking about how few things any of these CUIs need to know about, I’m optimistic that we can distill them down to a workable size while maintaining the LLM magic.
“Fridge, what is the meaning of life?”
‘Sorry, I don’t know about that. Ask me something about what’s in your fridge.’
“Okay how many eggs do I have.”
“I see 3 eggs.”
When I can have that conversation by proxy through my phone’s onboard CUI while at the store, I’m going to get a lot of value out of that.
So, appliances get even harder to understand settings, that are actually illogical, instead of just having hidden logic? That's not a clear win.
But transparently wrapped around that there’s a “good Clippy” who can teach, interpret, and orchestrate those settings with a CUI (conversational UI, pronounced “koo-ee”).
It is just completely against the modernly accepted "best practices" for devices and interface development. So I don't see how we can get it. But yeah, it could be good.
If it works like ChatGPT does, then I would find it a greater mental burden. You'd have to carefully craft what you're telling it, or engage in a conversation of some sort, instead of just hitting a couple of buttons or turning a dial.
If ML is a win for an IoT device, the hardware's been there for a while, though I'm sure yet-cheaper hardware might unlock a few more applications, it doesn't feel like much of a game changer.
The clear use-case for it is mechanical control and feedback. But it seems that every robot has enough tiny problems that you can justify a larger CPU anyway. I too am having a hard time being excited about it.
The potential is immense and exciting, but try as I might, I can’t seem to make even fairly simple things work reliably as I want them to.
One thing I love about embedded projects is that you can achieve pretty incredible reliability because everything can be so dialed in and isolated from points of failure in, say, an operating system. Trying to use ML for simple tasks felt like it eliminated that and introduced seemingly arbitrary failure into projects I really needed to work perfectly.
Again, just a hobbyist, so I can’t make any broad statements. It seems to align with what you’re saying though. Once we get past this hump and have more powerful/effective models at a reasonable price point, I feel like it could be transformative. At the moment it’s still extremely interesting and fun to experiment with.