Airbnb's preferred smart lock vendor accidentally bricks 500 door-locks
boingboing.net
boingboing.net
The scariest times were pushing updates that affected WiFi connection for customers with unusual network configurations. In particular, we had one with a multi-campus Cisco WiFi system with WPA-Enterprise, which requires logging into a RADIUS server to get WiFi. We were running on FreeBSD which had a half-assed port of the Linux WPA authenticator. We never managed to replicate their network perfectly for in-house testing, so pushing updates to them was always nerve-wracking.
Was this handled in userland? Or with something lower level (e.g., a hardware watchdog)?
I'm asking because the question has come up a few times recently about making FreeBSD recover (unattended) in the event that an unbootable kernel is installed.
sleep 1200
if ping -c 5 central-services.anybots.com; then
rm <this file>
else
cd /home/dist && git checkout STABLE && make install
reboot
fi
(We used git to distribute binaries, so we could checkout and install a previous version easily).This was a last ditch recovery. I could imagine lots of things that break, but which still allow ping. And installing a previous version over a newer version is flaky, because a newly created file won't be deleted.
An unbootable kernel is a whole different kettle of fish. I believe GRUB can be told to fallback to a safe mode.
I'd assume so; FreeBSD's loader can also do that. The problem was that if a kernel hanged during boot, we wouldn't get back to the loader until someone power cycled the system, so the "try the other kernel" logic wouldn't run.
After the upgrade, before the reboot, setup the HW watch dog for 2-5 minutes. Setup the boot-old-kernel flags as default in the uboot/EFI/grub/or your system boot loader.
If the upgrade the successful (connect to back to the server) and function correctly, disable the "boot-old-kernel" flags in your bootloader.
As always, do a lot of testing and automate those tests. You will be surprise how often a simple boot/reset/power cycle can fail if you repeat it over and over again in a weekend.
If you set up your watchdog correctly and the system loops too many times, your u-boot launch script can swap booting to safe/backup partition.
I can also confirm the wifi problem. Not knowing if a device is bricked or just on bad wifi is a huge problem you have to somehow account for. It was terrifying before I built tools to keep track of which devices those are
Just in case the domain name lapses. (Or for any of a hundred other reasons why having a system that reads plain text commands from the network is a bad idea)
Practically, there are tons of instances where executing code downloaded from URL's is how software is updated, but again most of those instances are human-driven, not automated.
If someone gains access to our server, its already game over. Being able to execute code on the devices is then moot because we were already hacked & potentially lost all credibility in the industry.
To everyone saying not to execute code at URL, this is how oh my zsh, and many other open source projects are installed. And SSL should prevent any kind of MITM attacks. The point about keeping keys off the server is a good idea, but is not always possible when the server needs to dynamically send content down to the device. The server would need to access the private key in order to sign this content dynamically. For statically built executables this is a good idea though.
1. A backup recovery read-only image which can flash everything back to square zero via some sort of hardware over-ride (like a pin-hole button that has to be held down for 10 seconds)
2. A temporary place in memory to buffer and validate the checksum of any new firmware that is uploaded before flashing it.
Real time flashing over spotty WiFi (without any sort of post-transfer checksum validation) has got to be the dumbest, laziest, cheapest, riskiest thing a hardware manufacturer can put out, and any manufacturer that does that should be ashamed of themselves.
I never think about my deadbolt unless I am opening my door. I never think about my smoke alarm unless starts to chirp because of a low battery. I never think about my light bulbs unless they burn out.
Any "smart" features, like notifications or a phone app, are going to make me spend more time thinking about these products than I care to, even if they are working perfectly. Failed OTA updates make things even worse. That door lock is going to need to introduce an extremely compelling new feature to convince me to spend time thinking about such a basic component of my home's infrastructure.
However, I've seen some airbnbs that have a pin door code or a key fob, which are arguably more secure than a regular key lock but have less chance of failure in my experience.
Many Airbnb owners do not.
https://www.techdirt.com/articles/20120823/10320820137/hotel...
For Airbnb, etc., you'd have to change the combination regularly (which is easily done in about 30 seconds) but at least I don't have to worry about getting locked out if I forget my phone, it's battery is drained, cell service is out, etc.
But I've been quite happy with my smart thermostat and sprinkler controller.
Maybe the difference is in these cases, a controller has been replaced with a smarter controller.
It's a natural line of progress, and there have been obvious deficiencies with the logic of the "dumb" controllers since their inception (heating the home when you're away, watering while it rains)
While software mistakes do happen, it seems any moron these days thinks they can put out an IoT product because they know how to flash a led on Arduino and they make no effort whatsoever in thinking how to make their devices have minimum reliability
Have a user accessible USB port (from the inside) from which a signed payload can be loaded for cases like these. Would have saved several days for their customers and several shipping fees.
Now that I think about it, I hope these devices are validating the certificate of publisher and the updates are signed.
You're joking right? Really, the vast majority of this iot stuff is ridiculously bad in this respect. Plain text data transfer, unsigned updates from any source based on a text message and so on. It's a disaster waiting to happen.
> from which a signed payload can be loaded
All updates should be digitally signed by the lock manufacturer
But then we complain because we don't "own" the hardware and can't hack around on it.
Also, no update should be carried out unless the user initiated it - for instance, a flashing led or an app notification tells the user an update is available - and it needs to be initiated by using the same controls as would be used for a roll-back, ensuring you do not end up locked out after updating.
That being said, the idea of a lock being online scares the bejeezus out of me. Call me a luddite, but as a minimum I want my burglars or vandals to show up, rather than nessing with me from wherever.
Well they can't be burglars if they never show up :p
In this scenario, the biggest problem is updating the bootloader, which cannot do this type of fallback.
Generally people avoid updating the bootloader (for good reason) but it is necessary sometimes...
Worst case scenario is if the system does come up enough for the hardware watchdog to get kicked in a timely manner, but nothing else works.
The watchdog needs to be the last initialized bit of code on the system.
Even so, determining when a partition is "bad" is hard.
An 8 MB (64 megabits) SPI flash chip[1] is about $2. Two jumpers to power the SPI flash and set the boot/load option is another 3 - 4 cents. These take up less than 30 sq mm of board space. And they give you a tremendous back up in case of failure.
[1] http://www.mouser.com/Semiconductors/Memory/Flash-Memory/_/N...
All of the ARM chips I'm playing with these days have a serial boot loader in ROM that you can get to by setting a pin on the CPU and powering it up that way. In such a scenario it would take no extra parts.
I believe they emulate or copy the system used by FPGAs to read a configuration from a dedicated SPI chip. Its well understood and a lot of tools can program the SPI chip to have the right sequence of bits in it to do the right thing.
If I were designing a system with OTA updates, then I would design something like this with it for this very reason.
These locks have some nice advantages but sometimes I feel like they are way too internet-reliant and a bit outside my control. Furthermore the company isn't exactly the picture of professionalism (their IOS app has actually gotten worse during the year+ I have owned the locks) and I wonder if there will be any functionality left in my $600 locks if/when they go out of business someday.
Just get normal locks and hire a damn property manager. When I used to work for an ISP in a resort town, our support desk would get tons of calls when power went out (a frequent issue there) or our service went out.
It was always property owners who lived multiple states away who couldn't get their renters access to their property. Always furious. Always without someone local who could let someone on to the property.
Could have been the Electronic Flight Bag, those can run commodity OSs. I can see an OTA failing there and it needing an OS reinstall. It'd be odd, but not impossible.
Nothing gets changed unless it's for maintenance or airworthiness reasons (or fuel economy)
Updates can be done during scheduled maintenance if there's a good reason for it (see above) or if it's really something urgent it is done overnight if possible or the plane gets scheduled for this specific maintenance
What will it take to get IoT device manufacturers to adopt development practices that prevent this kind of thing happening? It's bad enough that most software is bug-ridden and poorly tested, but it's inexcusable for IoT software if they ever expect IoT devices to replace traditional alternatives.
I wonder if a certification body (like UL) could work though. Probably is it is hard to boil "works good" down to a set of clear requirements.
I expect this to continue getting worse. At some point, hopefully, things will change and begin to improve. I think we'll see many more -- and worse -- incidents like this before that happens, though.
They do accept traditional keys but the whole point of the lock is to delegate restricted access to people at certain times and with logs/notifications. So a physical key is like an admin override, not to be used in normal circumstances. As such, only me and my partner have them, and we generally never use them. In a wonderful example of Murphy's Law we were both out of town this weekend so all the people that were supposed to be coming to the venue to set up for a show were locked out with no recourse, it was a huge mess. Long story short, the use case of these locks renders the traditional key fallback less than ideal.
Having to pay to send a locksmith to every lock they bricked.
Are you going to operate a crank to recharge the battery? I'm not too keen on wiring mains power up to something connected to my door handle.
This is a solved problem: https://en.wikipedia.org/wiki/Isolation_transformer
If the failure mode is potentially "you die", then we're probably better off low tech.
We had a similar debacle in early 2014 with automated updates. The crowdfunded Lockitron used a gear train that required it to back turn in order to unload the mechanism such that a user could turn the lock manually.
A hastily deployed update with insufficient QA would overturn the gear train past the limit of the underlying lock mechanism. At this point the stress would be released in one of three ways; (1) the main worm gear would pop off from the motor shaft, (2) a secondary gear would lose some teeth or, in a few rare cases, (3) the underlying lock would be damaged. We luckily had enough units to quickly ship replacements and rolled up an update that reversed the changes.
As a result, we made an explicit decision not to implement remote OTA firmware update capability for Lockitron Bolt. Instead we require users to be within Bluetooth range of the door so they can immediately verify the device is working as expected after an update. Given we don't want to nag users on a regular basis to update their lock, we opt for very conservative release cycles.
Additionally Lockitron Bolt uses a vastly simplified clutch mechanism that doesn't require any back turn behavior. This means you can always turn the lock manually (by hand or key), even if the device is dead.
I'm certain that these vendors don't have nearly the same amount of checks compared to smartphones due to much less overhead (like we had from carriers, Google, SoCs), but it's still astonishing how this could happen, but it's a good lesson for them, mistakes happen and I hope they implement a system where this is not possible to happen again.
The solution is simple, though: Use U-boot, two partitions for firmware and a tiny script using a boot counter/flag. Only after successful boot set the flag "do not boot back to old firmware". But doing that well requires recent versions of U-boot, and people are STILL shipping devices with 2012 or earlier versions...
Smart implies some sort of intuitiveness or responsiveness, which is often not the case.
LockState makes the mistake, and AirBnB takes collateral damage.
If a system can go back the last working version if the latest fails to work, even a bad version shouldn't brick the device.
Doubtful, dead simple OTA with recovery requires a 2x allocation of space. $2 or so of parts covers it for 99% of embedded scenarios, and if the product doesn't have enough storage to do a recovery OTA then the product wasn't spec'd properly. More clever but just as robust systems can get away with less (however much a recovery partition takes up, if doing WiFi that can be a lot).
> If a system can go back the last working version if the latest fails to work, even a bad version shouldn't brick the device.
Knowing when the current version has failed can be hard. If the system comes up 99% of the way but one driver fails, that can be hard to determine programmatically. Obviously the manufacturer here didn't have that robust of tests running at startup to determine if there was a need to rollback.
To be fair, most software updates don't have that level of rigor. Pre-release testing catches the majority of issues, as it should. Now days, phased rollouts to opt-in beta testers is standard to try and catch any remaining issues.