Learning bimanual mobile manipulation with low-cost whole-body teleoperation
mobile-aloha.github.io
mobile-aloha.github.io
Imagine a high-end apartment or condo building. In the basement of the building, there is a fully equipped professional kitchen - however, it's not supposed to be used by humans. Instead, the kitchen is fully automated with smart/wired-up devices wherever possible. The actual chopping, cooking etc is performed by robotic arms like the ones in the OP. The kitchen has a human-accessible storeroom on one end (that must be regularly restocked by the condo administration) and is connected to the apartments by a system of service lifts.
Residents can order food through a an appliance or an app or something and choose from some menu of predefined meals (or possibly even create their own recipes through some sort of building block system). Once they submitted the order (and all ingredients are available), the kitchen springs into action, prepares the dish and places it in a service elevator. In the apartment, the dish appears, fresh and hot, like right out of a Star Trek replicator.
I guess we've just gotten one step closer to that vision :)
It’s also amazing that there’s just instructions and it’s open to anyone. Kudos to the builders. They sauté shrimp, do laundry, and wash dishes with this robot you can build at home.
If that isn’t cool I don’t know what is.
Having to do it everyday for years and not being able to do other things because of that is not fun at all.
Jokes aside -- something like this could be a fantastic breakthrough for people who've lost the use of their limbs. A wearable / chair-attached exo, driven to do complex tasks through simple commands.
If I understand it correctly, it learns from "watching" you during telepresence sessions? If so, and if it learns fast enough, it really would be a massive boon to some disabled people.
Although I want to see it peel potatoes and dice onions :)
The "rinse pan" demo did not impress - but it doesn't need to. You don't have to teach it to "scrub bowls and plates" when you can teach it to "fill and empty the dishwasher"
As an outsider with no knowledge of robotics, I've always been surprised that manipulation tasks (or just smooth robotic movement) are so challenging and seem to progress so slowly, especially when compared to IT.
The amount of feedback built into your end-effectors (pedantic word for human hands) is insane. If you're not familiar, proprioception is a good google/wiki hole. Most of the signals that allow you to move your hands don't even hit the brain stem, let alone the boss upstairs.
The challenge mostly lies in how we've instrumented these things. Precision requires low tolerances. Low tolerances + unexpected environment == you've just driven your robot through the countertop/pan/coworker or broken a very nicely geared servo.
These days we have 3d cameras, but they still only see part of the objects we want to manipulate. The back side is hidden. So you need to either specify and model all objects to interact with, or have some word of a world model where we can predict what the full object, it's weight, center of gravity, surface texture, etc, is like.
And before we even decide to manipulate it, we have to detect it, categorize it and segment it (where does the pan stop and the stove begin?). We have to plan out a manipulation task, including finding grasp points, finding movement patterns that do not interfere with the rest of the environment, etc.
It's a whole bunch of separate problems that need solving all at once. There's motor control, building the right manipulators with the right sensors, bringing all the sensor data into something where we can make a single decision, understanding of the world and what happens during manipulation, and higher level planning.
I'm far removed from this field, and speaking as a layperson, so pardon my ignorance.
We don't even know if "intuition" would arise from the knowledge you claim, we don't know how that model would work, and even before that, collecting all the data (not to speak of availability of all the sensors) is a vastly more complex than even what ChatGPT or any LLM model data collection would ever be.
>it's own experience from reinforcement learning
This is a common mistake often heard from CS -> ML(RL) -> robotics transition folks. Reward function is given for free in RL, but in the real world, estimating the reward is a complex problem in its self. That's why RL on robotics have mostly seen success in quadrupedal locomotion; the reward function is simple (forward velocity, calculated from IMU), but how would you calculate a reward function in 30Hz+ for a simple task such as "chop onion and put it in the pan"? If you can construct the reward function for that task, most likely, you already have all the world-states available and might as well skip RL and do something else with that, such as Model-predictive control.
As for intuition, see: https://en.wikipedia.org/wiki/Moravec%27s_paradox
I love the quote at the end of that article you linked:
> As the new generation of intelligent devices appears, it will be the stock analysts and petrochemical engineers and parole board members who are in danger of being replaced by machines. The gardeners, receptionists, and cooks are secure in their jobs for decades to come.
I should've picked a safer career in gardening...
This other content really jumps out at me as well because it's extremely true.
Even older than walking and manual dexterity are really basic abilities like eating. We're nowhere close on that - were so far off it's not on anyone's radar. Robots will run on batteries or some other form of power - there is no way anyone is close to building robots that can eat break down food and use it for energy and repair. One of the oldest evolutionary traits.
The other is course being procreation. Will a robot be able to assemble a new one from pre-made parts? Likely not too far off. But could a robot build or grow one from scratch? That's so far off in the sci-fi future it's silly.
Couldn't we sidestep the complexity of digestion and just get energy from the Sun? With improvements in solar cells and battery technology, we wouldn't need to engineer something as complex as extracting nutrients from food.
I don't think we'd want to replicate biological systems in robots. Digestion and procreation happen at the cellular level, and achieving that with technology is indeed hard sci-fi. Autonomous humanoid robots can exist and be useful for us without this level of sophistication. Though once this happens AI itself will be capable of self-improvement, so we can leave it up to them how they want to improve. I, for one, welcome our new robot overlords. :)
I think one of the issues is that in some parts of academia, progress is made one PhD at a time. And a PhD is almost always too narrow to bring all of these fields together. I'm sure they are solvable problems, and I'm sure they will be solved. But maybe it will take some other research structure? Private? Guaranteed long time funding for academic teams?
So yeah, smooth motion feels easy, but is a gd miracle of biology :)
* Robots can make up for a lack of prediction through really really fast control. This is how Boston Dynamics robots operate at a basic level.
could you elaborate on the mechanical/physical limitations that cause SOTA actuators to lag behind muscles, and if there's an equivalent "moore's law" that might predict when this gap closes appreciably, if ever?
No doubt there is lots of much more advanced reasons. These are just from the top of my head.
Basically, such a motor can lift two chocolate bars (2x100g) at arms length.
https://twitter.com/tonyzzhao/status/1743378437174366715
In another tweet [1] the authors give a count of the successful (but not the failed) attempts at various tasks:
Our robot can consistently handle these tasks, succeeding:
- 9 times in a row for Wipe Wine
- 5 times for Call Elevator
- robust against distractors for Use Cabinet
- extrapolate to chairs unseen during training
https://twitter.com/zipengfu/status/1742602883256943040___________________
[1] How do you call those now? Exes?
All those are toy tasks that have no application in a real environment, even the constrained and save environment of a home. Most homes don't have such large empty spaces. Most of the time when you need to tidy up a bunch of chairs they're not put in a neat straight line by a RA, they're left in a jumbled mess by a stampede of students and you have to do a lot more pushing and pulling and turning around (and there's tables and possibly empty cups and stuff). Most of the time you need to cook a lot more than one measly prawn. And let's not talk about the primordial chaos of kitchen cupboards.
What's worse with all those demonstrations: the robots can only do exactly what you see in the video. Change the parameters even slightly: different shape pan, different height cupboard, different room configuration; and the magick -poof- vanishes, into thin air.
That stuff doesn't work. We aren't even close to solving autonomous robotic behaviour. RL doesn't work and the older techniques aren't working either (planning). All that stuff ever does is get published into papers, advertised with fanfare and then forgotten because it never makes it to the real world, because it's all unreliable and unpredictable and costs too much if you want to do anything real, and that's always something trivial. So the state-of-the-art in robotics is in hand-coded industrial robots that do one thing and do it over and over again and nobody asks them to generalise, or to come into your house and cook you a prawn.
That (hierarchical task planning) is a problem that is not yet solved (because it's combinatorially hard) and eliminates any idea of "general" ability: you have to set things up perfectly before the robot can do anything. Exactly like for industrial robots. It just looks more general because it's inside a house rather than a factory.
What's wrong with only being trained to use a specific tool? The robot is designed to be easy to train via teleoperation. If you told me that one day it's going to be cheaper to program a robot to press a button than to remodel the switch to be robot friendly I would take these demonstrations as evidence.
The things you are complaining about are quite trivial. For example, they could have added a second prawn...
Edit: I just realized that it is only the first video was sped up by 6x. So basically, the robot is already fast enough and can only get faster.
What's wrong is the fact that it's not just one tool, but one particular tool (specific size, shape, colour etc), placed in one specific way on one specific countertop, in one specific kitchen etc etc. Try to calculate how many different configurations there can be of this combination and you'll convince yourself that there is no way to make any real progress with methods that rely on such specific conditions and have to be re-trained everytime any detail changes.
As a for instance, the robot trained to fry one measly little prawn would have to be trained again to fry two, or to fry them in a bigger or smaller pan, etc.
Btw, the robot didn't fry the prawn. It pretended to. Because it has no way to tell when the prawn is fried. Good luck training it to do that.
>> Edit: I just realized that it is only the first video was sped up by 6x. So basically, the robot is already fast enough and can only get faster.
I was talking about the initial video, not the others. And you're wrong- that's the limit of the hardware. To make it faster the researchers will have to build a faster robot and that still doesn't change what it can and cannot do, at any speed, fast or slow. The speed doesn't have anything to do with generality.
>> One day it will be 3x slower, then 2x, then 1x, then, oops.
One day we will all live on Mars and consort with aliens. I like science fiction too, but this is proposed as real-world progress, right now. Well, it ain't.