Gemini Robotics 2 brings whole body intelligence to robots
deepmind.google
deepmind.google
Isn't that Checkers right from the dawn of reinforcement learning?
https://www.engadget.com/2225849/google-shuts-down-alphafold...
Google acting like a normal but competent company that's just chugging away, meanwhile the hype cycle is propping up its upstart competitors to truly ludicrous valuations.
Chugging away not releasing products or releasing ones worse than their competition typically. Not sure that’s what a normal competent company should be doing.
Competent is a stretch. Google's AI offering seems to be, once again, PM led–lots of constantly-changing brands being merged and deprecated with zero customer support or service.
They'll almost certainly be one of the survivors. Their technical competece is unmatched. But that doesn't mean they have a great product in the way both OpenAI and Anthropic do. (Outside their datacentres, which are a legitimate feat.)
For a company that prides themselves on the smartest engineers, they really struggle with the basics
> Google acting like a normal but competent company
Is it messed up and weird? Yup. Still normal though. Worse, this is matching the baseline of competence I find in so many other places.
Furthermore, Gemini is just all around a less trustworthy and mature model, for many reasons. Very smart but lacking the precision and holistically exhibited in the more recent models from OpenAI and Anthropic. On the flip side, Google's work on Gemma is unmatched.
The thing that gets me with Google is that they've killed too many products I used daily (RIP Inbox and Google Podcasts) for me to be able to trust that any new offering will exist in 5 years.
“I also see that I have two Microsoft Outlooks, and neither one of those are working if you want to remote in and check"
Cdr. Reid Wiseman to Mission Control, Artemis II, Earth-Moon Coast.People are too quick to forget racially-diverse Nazi soldiers and other hilariously incompetent stuff that has been going around their generative AI efforts just a couple of years ago. Google look like they finally getting their shit together under immense competitive pressure but I would keep my eye on them for a few more years before starting calling them "competent".
Didn't all those authors leave Google?
I am just waiting for the 10 AI models or Chatbot Apps competing with one another like the 10 messenger Apps.
Gemini
Gemini Live
Gemini Code Assist
Gemini for Google Cloud
NotebookLM
Gmail
Google Docs
Google Sheets
Google Slides
Google Drive
Google Chat
Google Meet
Google Calendar
Google Forms
Google Vids
AppSheet
Workspace Studio
Google Search
Google app
Google Chrome
Google Lens
Google Maps
Google Flights
Google Hotels
Google Shopping
Google Photos
YouTube
Google Messages
Google Translate
Android
Circle to Search
Google Home
Google Play
Pixel apps and features
Google Flow
Google AI StudioI think LLMs might have other use cases, but all the ones that were dreamt up basically disappeared.
Google is losing in the race that really matters
I don't think we'll get household robots anytime soon. Hell, the only ones that can afford them will be the same that would hire human household help.
But maybe they'd rather deal with a robot than a human?
I do want a robot serviant. But running AI that I control. Here it google says it can "And this profound intelligence can also run locally on-device", that makes it interesting for me.
I don't want a person I barely know regularly going all over my apartment. What if they discover... nevermind.
A robot, on the other hand? Hell yeah!
My issue is different: I feel bad with someone working around me while I am lazy. A robot would not pose that problem.
That aside: if I'm lucky enough to find a person who's good and reliable, they might move away, switch jobs etc.
It has all the headaches that come with hiring and managing someone, because, well, it is exactly that... If I don't want to be a manager at work (been there done that, happy to let others do it and get the raise that comes along with it), I sure as heck don't want to do it at home.
Those controlled by LLM's maybe ain't.
Here with google the LLM is even local, but will be fun with the cloud controlled robots, where the company got overwhelmed and silently uses a weaker model at certain times. With fun results.
In the humanoid robot scenario, you'll get a surveillance device with a built-in microphone and camera beaming every intimate detail of your space to the highest bidder. Instead of getting one person you don't know very well in your space, you'll get thousands.
Just imagine, there’s a woman Alice who cleans your home once a month. On this particular day, she had a fight with her husband this morning on (something completely unrelated to you).
Do you expect her to check her emotions at the door when she enters your space? Or is she going to give really bad vibes while she is moving around your house?
This is the same reason people prefer buying Tesla cars direct instead of having to make a deal at a dealership. If they all had high EQ maybe they'd prefer the dealership model after all.
Tech does also attract grumpy people.
The reason this idea stuck in my memory when I read it - is that it resonated with me. To some extent I feel this way.
You have cause and effect wrong; it just makes people that have to deal with the people in and around the tech industry grumpy.
More seriously:
- I wouldn't feel obligated to 'pre-clean' anything. Who wants to come across as a thoughtless/careless slob whose personal habits amount to borderline-abusive demands on the cleaning staff? Not me... and where does that line get drawn, anyway?
- I wouldn't worry about stuff being picked up and misplaced
- I have never had problems with maids stealing stuff, but I know others have
- Cleaning can be initiated or postponed as needed, with no dependencies on someone else's schedule
Well, as long as you don't mind the workers in the robot company's telemetrics department whacking it to your robot's video feed.
(I'm being facetious here: nothing that transmits data out of the house is going to be deployed to run household robots, at least not in my house. Even if I didn't care, I can't make that call on behalf of family members or guests. Then there's the obvious concern that the robotics company would sell the data to my homeowners' insurance company, health insurance company, and who knows who else.)
That being said, by the time the kinematic hardware needed to build a robotic housekeeper is available for home use, I don't think the brainpower is going to be a problem. There's nothing in Ex Machina that couldn't run on a couple of RTX 6000s.
(real-time mess handling)
For example, you could store far more things out of easy human reach.
For businesses the bar for adoption is very low: If the thing can work repetitive jobs for 24 hours a day and replace 3 shifts, the purchase bar is nominally anything less than 3 x human salary if your budgeting horizon is 1 year. That's a high number, and probably fairly easy to achieve.
For homes, it's a very different bar. You'd have a hard time convincing most American families to purchase anything with a >$1000 price tag. Currently that's pretty much impossible for a humanoid.
Like cars, they'll likely purchase them through finance deals, with a monthly payment. Or perhaps even rent/lease them.
Robots aren't a necessity, and for a price tag in the tens of thousands, most people will just mop their own floors and do their own laundry.
And (regardless of the costs involved) if I were going to pay money to not do this stuff myself, I'd rather pay a human to do it because I know there are humans out there that could use the work and I value them much more than I value Google and other corporations making even more obscene amounts of money selling future e-waste.
increasing your spend by 25% is a smaller change then increasing your spend by infinity%
If Rosie was the only output she would never happen. Rosie might happen as a side benefit of replacing American workers but only for the rich people.
The vast majority which may be replaced would probably be better off burning down the factory rather than celebrate their upcoming domestic helper.
A human house cleaner by definition is a real human looking at your home every time.
Versus a small army of roomba's and equivalents mapping out floor plans and tangentially exposing human occupation times with that large data set almost certainly exposed to criminal and surveillance networks that will almost certainly use some of that choicer data to their advantage.
It's literally trade offs in both cases.
But that means speed does matter after all. The renting company will need to clean as many flats as possible with a single robot in a single day.
So the lifestyle stayed inflated.
1. 80%+ of Americans can't afford house cleaning services. I'm a software engineer in silicon valley and even I clean my own house because I don't think I can afford it. My rents have gone up by more than 5K in the last year, my salary hasn't, so that's coming out of my house cleaning budget.
2. Even if you can afford 5K/year on it, the cleaning service is not an upfront commitment. You can bail anytime when you get laid off and clean your own house.
3. Consumers don't budget rationally. Businesses do.
As opposed to a household which might not have enough work to fully utilize it, or just have trouble scheduling times where it can work without causing disruption.
How big is your house? When we had a cleaning lady it was $70 a go. Under $2000 for every second week.
Wondering if you're closer to the southern border or something.
I think I remember maybe around $100 including laundry almost 20 years ago and that was a cash type arrangement, weekly including laundry (multi hour visit).
You can pay over $500 for an automatic cat litter box. Robot vacuums can be cheap but run to $600 or so. Household appliances is a robust sector where people have a proven track record of spending a lot of money.
This is my dream for a house chore robot. If I could dump hampers of clothes into a receptacle and get stacks of folded clothes out, I’d happily pay $1k. Household of 5; my kids each produce 3-4 sets of clothes a day (sleep, school, sports, after school). I run wash in AM so dryer finishes before 3pm (pg&e ToU) and then I’m focused on other things for rest of day. More often than not, I get to bed to find a near full hamper of clothes dumped where I sleep and then have to sort/fold/deliver. Bonus points if robot can sort different items to different stacks/bundles.
I don’t need AI/robotics to save me time from having to think, research, or code… I need AI/robotics to save me time so I can think, research, or code.
Interested. Which one?
Great machines!
If I can get a robot replace a gardener who I pay $170/mo for 2 visits per month, i probably would. But I suspect that would kill lawn
Apparently prices on them have gone way up. The one he has now would cost something like $6K today.
> You'd have a hard time convincing most American families to purchase anything with a >$1000 price tag...
Which I am skeptical of.
I think house cleaning/chores require at least a "part-time" job's work of work for the average household (especially with children). We may be atypical, but my partner and I don't want to hire out a cleaning service and deal with the whole human component. But, I think we'd gladly pay 15k+ if it allowed us to take on the additional gainful workload. Even at a 20k price tag, I suspect we would likely have made up the difference in under a year.
Selling to businesses is much easier, 99% of businesses can afford 200K if it replaces 3 x 70K humans.
I'd gladly pay or finance a robot that could have dishes clean, laundry folded, and carpets vacuumed by the time our family got home.
Mileage varies. My impression is that on net, many people borrow money for the vehicle and end up adjusting their lives to depend on it, so much so that in their head they think they're saving time, but if they'd put that money into their home budget and moved closer to where things are, they'd get more exercise, be more socially healthy, and also waste less time and money driving. In aggregate, this makes everything worse and perpetuates the cycle. I sure as hell wouldn't move somewhere I had to drive to and from if I didn't already have the car, and I sure as hell wouldn't buy a car to get to a job unless the nature of my work required it. Not again.
They only became a “necessity” once they went mainstream and the world started evolving with vehicles becoming a part of everyday life.
I can certainly see the same thing happening with residential robotics.
Or, if the stress of doing everything they do is too much, and interferes with your ability to work and do other things you once had the time to do.
I've used on in the past, but I found that I was routinely dissatisfied with how they were doing things and I felt I needed to maintain a certain level of tidiness to allow them to clean properly. I think there is an intrinsic cost to having hired staff in one's home that most consumers would be factoring in, consciously or not.
Probably $80/day for someone to do everything I want done. 20 days a month.
$1,600/month.
It'd take two years, and the robot has paid for itself.
Can a household robot provide as much utility as a car? If "yes", the market gets pretty big.
Robots may not be very useful for now, but they will be at some point, and people will buy them, financing if they have to. Even if what the robots can do are not non-negotiable.
I pay $200 every 2 weeks to have a 3000 sq ft house cleaned.
That's about $5200 a year to fully clean a house. And it's a human, so they can also tidy up, clean out the refrigerator, do my laundry, water my plants, etc.
My parents, aunt/uncle, grandparents, and one of my neighbors uses the same house cleaner.
If she stole from me, I’d immediately tell my family and she would lose basically all of her clients.
That makes it easier for me to trust her.
(It works similarly in NYC apartment buildings - a single cleaner often services many people in the same building, so they have a reputation to uphold)
I would never use random online cleaning services.
I'm sure that there are services that do this, but I very much doubt they are the $400 per month services you are talking about. Robots may not be able to do that level of task at all right now, but they will get there.
And of course, that's not even thinking about all the cleaning tasks that would be avoided in between the every-other-week visits.
Having my surfaces be somewhat cleaner every two weeks is worth ~nothing to me. It's by far the easiest and least unpleasant part of cleaning a house. Having a robot that will pick up the detritus of everyday life and put it away, pick up and clean dishes, etc. and do so every single day would be worth an enormous amount of money.
I have no idea how long it will take for robots to get to that point, but whichever company cracks that level first is going to make a lot of money.
And of course, that's completely ignoring the fact that I also would like to not have to pick up my shit if I could buy a robot that would do it. Yes, I can. And in fact I currently do...mostly. But I'm not happy about it and if there was an even quasi reasonable amount of money I could pay that would mean I would never have to do it again, I 100% would do so.
And if you can honestly say that you wouldn't...well I hope you are enjoying your life in the monastery.
2 year olds do a bad job. They are 2. Yet it is valuable for them to learn to manage themselves. They love helping put silverware away from the washing machine. Kids _need_ to help as part of their development.
I won't lie, I don't do this often, and I end up picking up plenty of toys, and would welcome robot assistance, but it still seems wasteful and privileged to use advanced tech in this way
A new category of task starting to move along this trajectory is to be celebrated. Maybe some day people will clean some small portion of their house the same way that some people now have backyard gardens.
May be they will. But do you think this because of the LLMs? I don't think LLMs will help. Because it can only generate text.
Costs me 160 / month for 4hrs / week.
Obviously the whole question is how a robot cleaner compares to a human.
As you point out, the potential for robots far outperforms humans in multiple categories (although clearly not all by any stretch).
I'd say pricing around that of a lower end new car - or nice used one - seems about right to me. In terms of justifying the price, another thing I haven't seen discussed much is the competitive advantage this can provide. If you're spending less time cleaning, that's more time for other things.
Some people might feel that a robot cleaner is less intrusive that a person.
People might not have 20k in the bank, but if most are doing a part-time employment's worth of work which could be replaced by a machine, it would make sense on paper.
Granted, people aren't rational. Anecdotally, my partner and I probably could be hiring out some of this currently, but we choose not to as we don't want to have to work around a cleaning service's schedule. I think we'd be more apt to buy something which we own, have control over, always does things in the manner we'd like, doesn't have to work around our messy lives/schedules, etc.
I could definitely see people baking in money for certain automations, if they have a comparable ROI.
Of course, the initial prices will be significantly higher than "end result" or commoditized prices, at which point maybe even people on the brink of poverty can afford some level of home automation. In the interim, yes - it will be largely for the wealthy and business applications.
I kind of understand where the feeling comes from, but this also saddens me. What happened to our generation? We're scared of people.
I share those feelings too, it's self-reflective.
If the future home robots are any good, it saves you from buying a dishwasher and robot vacuum. It can replace a maid/cleaning service and gardener/lawn-mowing service. For anyone paying for those services, it pays for itself.
$1000 isn't that high. Probably about a third of American households have an appliance over $2000, which is still a massive market.
The side benefit is all your dishes are always available. With the dishwasher, up to one full dishwasher load are dirty at any time, which means you need more dishes and more cupboard space than someone without a dishwasher.
If the robot is going to clear your dishes and wash them for you right away, what's the dishwasher for? To save the robot 10 minutes?
I'm not saying people who already have a dishwasher will throw it away. But a robot makes installing a new dishwasher much less tempting.
At my utility prices, it costs about 5¢ to hand-wash one serving (plate, pan, glass, utensils, and a bowl) with running water. That's not nothing, but it's dwarfed by the energy use of using the oven and the supply-chain energy and water use of the food if it includes meat.
People will casually waste 10x that energy without a second thought when they don't need to just because it's convenient, or because they're used to doing things that way.
The dishwasher looks more unfavorable if you're considering installing one in a small apartment or condo. Apart from the cost of the dishwasher, you need space to install it, and more cupboard space to keep more dishes, glasses, etc. For an extra 4¢ per person-meal, having your robot do the dishes buys you the convenience of skipping all that and having all your dishes clean and ready all the time.
Honestly I don't even bother with two cycles and just use a 45 minute program in a 45cm wide dishwasher while throwing the extra powder in there anyway because it is too cheap to care about the waste vs tabs.
And talking of kids, these will have to be exceedingly safe, a 6ft machine falling on a child is probably a worry most households could do without.
They all have cars. Turns out you can finance things, and turn them into forever debts. Paying $30K to never have to do chores again, or maybe even cook again, and live in a clean neat environment is not only an amazing proposition that opens up a lot of free time for people, but will also have a lot of social pressure behind it. We have social pressure that causes people to buy >$1000 phones to virtually no marginal benefit over $20 phones.
They will be paid off at $500/mo over 10 years, and most renters might just use the one that came with their apartment. If they're more like $80K or $100K I could see having a problem selling them. But if they basically turned apartment buildings into hotels, they would be a bargain for landlords; just give them the keys.
edit: of course, sci-fi has already rehearsed this. You can watch any number of movies and tv shows with families walking through a robot showroom guided by a guy in a cheap suit offering to give them the best deal. I'm sure you could find written examples from the 40s.
There are plenty of white collar employees in cities paying $500+/month for someone to come clean their home. As one of those people, I could easily see paying low five figures for something like this if it actually worked well and could replace most household labor.
Part of the reason is upfront commitment is scary and risky for personal money, but not scary for a business.
As with most things, it'll cater to that demographic and then filter down.
what you have to keep in mind is homeowners salary is way above avg salary so anything you can do to effect them has a much larger impact than it would seem at first.
Having said that, it could still be decades till the robot is ready to be released into human homes unsupervised. The variation in human homes, what people store in them etc. is not easy. However, it could be ready for hotels much sooner.
it depends. in order for my slow ass Roomba to clean my floors, I have to move a bunch of things out of the way and not use the room.
on the flip side, i think this "the robot cannot automate the whole task" thing is reductive too. suffice it to say, EVERYTHING matters, there honestly nothing that "honestly doesn't matter."
There is going to be so much pain from VLA malicious compliance. If the current gen of LLMs are anything to go by I can already see it being an hilariously massive problem. People are careless.
I only wish I could convince some humans in my life to slow down and do things with care.
The difference between gpt 2 and 3 was insane. 2 could generate limericks when it wasn't repeating a word 300x. 3 could actually do some things. By comparison gemini robotics has hardly changed at all.
I will also point out that slow, non-fluid robotics is on a totally different level of difficulty from fast fluid motion. Asimov could walk pretty smoothly, but it didn't fall over because it used a very careful sequence that was never unbalanced; you could pause at any point without falling over. Move faster, like boston dynamics, and you need to account for the change in balance from your arms swinging... or rather, you need to be able to account for the rotational inertia etc from moving multiple masses along complex paths with multiple points of articulation at hundreds or thousands of times per second.
An algorithm to fold tshirts 90% of the time is easy. The cloth hangs down by gravity and you can just look for right angles (corners), find their coordinates with binocular matching, and move them to meet each other. Getting 99%, or folding them quickly, so that the fabric is actually moving instead of just hanging still- incredibly, incredibly more complex.
TBF fabric is much more difficult to simulate than a bipedal body is.
As I recall openai had mujoco playing soccer nearly 10 years ago. Obviously real world bodies are much more difficult but I'd be curious to learn why that is.
Then there's the problem of generalisation. RL is insane in figuring out the dynamics of any arbitrary environment you drop it in at any time. Unfortunately once you have a policy trained in one environment (or one task) you have to train again for the next environment (or task) you want to deal with. The number of different environments, situations or tasks in the real world ends up being overwhelmingly large.
So for example DeepMind has been trying for ever to train various robots to do stuff like grasp arbitrary objects by training in the real world or in simulation and still grasping is an unsolved problem, in the general case; and in fact in most cases.
With LLMs it was different in that there was a www of text to train on and the stakes are lower because there's no sim-2-real gap and there's no need to deal with the real world, in training or deployment, it's just text on a screen. That's very hard to replicate in robotics.
There are major breakthroughs still to be made and they aren't currently even on the horizon. And no, I don't have the answer with my super secrete robotic AGI homebrew project either, just saying :P
A waymo is basically an industrial strength robot. It operates among inherently unpredictable humans already. There is no guarantee that it will never, ever hurt a human.
And yet in many cities around the world, you can call a waymo and ride it and society accepts it.
Is this the new "x technology is a year away" ?
LLMs aren't themselves hurting anyone, no matter how much certain people like to pretend otherwise, whereas an AI in a robot with significant motors in absolutely can and will. There are reasons industrial ones live in safety cages after all.
Expecting them to be like inverse kinematic driven digital dolls is wrong because the optimization won't be for matching that but something like "net reduce energy consumption" which for electric motors in multi joint arms will look a bit odd.
I did professionally few prototypes with robots and progress is real yet very far from what the average customer would find reliably useful in menial tasks.
FWIW I do think https://rodneybrooks.com/why-todays-humanoids-wont-learn-dex... remains relevant, namely dexterity is also a hardware problem, grippers aren't hands. They even clarify "multi-finger dexterous manipulation remains challenging." and those aren't even fingers with a lot of sensors.
It's certainly not there yet for anything practical, there's also certain bits and structures that don't have accurate names during construction, and it is important to keep that in mind - so a robot is unlikely to understand what it means to say "put the left bit of this box onto this right bit" due to ambiguity, a human would understand that
Plus we have no good reliable accuracy testing data in most cases (most tests occur on a few demos, but that isn't a good representation of how must things work), popular benchmarks, such as libero have been saturated, and nearly everything gets 95% there, most companies and researchers have their own benchmarks here.
Plus companies lie alot, and do very dangerous things in thier videos, I.e. these robots should not be standing very close to humans, because of being dangerous.
There are also legitimate concerns of misuse of these robots that need to be accounted for, misuse does not have to be warfare, but can be as simple as confusing it while it is cutting tomatoes with a knife.
Turning doorknob is easy, and fail recovery is also being worked on, but we don't have reliable statistics anywhere on that. The hard part is on practical things, as in when placing bricks or attaching a part during manufacturing it needs to ensure that it is aligning everything correctly....and that's hard, while it is impressive, it is very irresponsible to keep humanoids at home (people are irresponsible when untrained), for example, lawnmowers injure about 6400 people a year...and that is not an everything machine.
Humanoids in general are...not appealing in specific, due to maintainable of joints, complexity, but robot arms in particular, expecially on wheels (check mobile aloha), are likely to be able to do tasks such as clean up in hotels, after a guest had left, or replace some cooks in restaurants (if their work is consistent)
It's difficult to grok this because if you watch a human folding a t-shirt, you can reliably predict that the same human will fold a different t-shirt just as well, and in fact be perfectly capable of folding a wide variety of other clothes items as well. Not so for robots. With robots, what you see is precisely what you get. If you see a robot folding a t-shirt, all that means is that that particular robot can fold that particular t-shirt. The state of the art today is that the same robot cannot be expected to be able to fold a different t-shirt.
For example, see this article about Mobile ALOHA at Google. There's a passage where the visiting, awe-struck, journalist asks whether the robot he's just seen tying up a pair of shoelaces can tie up his own shoe.
“If I gave it my shoe,” I ventured, “would it just totally fail?”
“We could try,” Tompson said. I removed my right sneaker, with apologies to anyone forced to handle it. Tompson gamely placed it on the table, while Driess reloaded the policy.
“To set expectations,” Driess said, “this is a task that is thought of as being impossible.”
Tompson eyed his new experimental subject with some trepidation. “Very short shoelaces,” he said.
The policy booted up, and the claws set to work. This time, they poked at the shoelace without getting a grip. “Do you give consent for your shoe to be destroyed?” Driess joked, as the hands grabbed at the tongue. Tompson let them try for a few more seconds before hitting the Failure pedal.
https://archive.ph/CiJJG#selection-1887.0-1915.291
(Original: https://www.newyorker.com/magazine/2024/12/02/a-revolution-i...)
As to ACT-2 which basically uses the same techniques as ALOHA (imitation learning) far as I can tell, that's a commercial product and the information they give on their site is difficult to parse. E.g. they say they have 99.1% ±0.3 success rate, 778 successful folds and 9 garment types which is low enough to engender some trust they're not trying to inflate their numbers, but they don't say whether they trained on the garments used in evaluation or not. Chances are they did, because that's the current limit of the technology, i.e. if the garment being folded is unseen (as opposed to the environment, which they tout) then performance is basically random. So either ACT-2 have a major breakthrough that is a few leaps and bounds away from the current state of the art, or you've just watched a tech demo.
The fact that they only advertise "9 garment types" though is a big hint: they have the same problem with generalisation as everybody else at this point in time.
In the left half of the video you can see that the robot is (trying to) exactly match the folds of the human in the right half and it's doing so while folding the exact same garments on the exact same surface.
The right half of the video is not a training demonstration, I don't think, since the robot is trained by teleoperation AFAICT (it needs to because it must use its head-mounted camera to control its movements) but that just underlines the degree to which their training regime is exactly copying the movements of a trainer, on the same garment, in the same environment.
This is a limitation of the training approach, by RL. With RL you learn a mapping between sets of pixels (as in the video that comes in through the robot's camera) and robot actions (as in actuator commands). What that means is that once a policy is trained and the robot is deployed, if the input pixels are significantly different than the input pixels at training the robot doesn't have a policy that matches the input pixels and so it can't find the right actions to take. So they have to keep the training and deployment garments and even the environments the same, or as similar as possible.
You can see some more evidence of this in the video right under the paragraph with the title "Hill-Climbing Reliability Through Post-Training". The robot at the front of the video, with the bright red trim, is shown trying to fold a grey t-shirt with white flower decorations and a frilly hem (how adorable <3). But you can tell it fails because the video stops before the robot has completed the fold. If it could complete it, you can rest assured that the video would be showing off the entire folding sequence as it does for the robot with the green cap on the other side of the bed.
That paragraph is making a claim about a "post-training" regime that's supposed to improve generalisation but it leaves more details to a "separate technical post". So I can't tell what it's supposed to be doing, but I don't think it's working.
When I watch videos like that I always remind myself that a) I'm watching a tech demo created to attract investment and b) I've watched way too many of those, going all the way back to the Boston Dynamic videos of Robot Dog or of Atlas doing backflips and yet the state of the art hasn't really budged since. Such videos make it easy to overestimate the state of the art in autonomous robotics and in fact are meant do precisely that: play up robots' true capabilities. It's just impossible to say anything about a robot's general capabilities by watching a few minutes or even a few hours of video. OtoH if you know what to look for you can tell everyone is basically stuck at the same level and trying the same things to escape it. The truth is robotic autonomy is several major breakthroughs away and nobody has any idea how to get there. So we'll be seeing many more of those tech demo videos in the years to come.
Humanoids are far from being remotely useful in real life scenarios, yet.
Fun fact: most (if not all) qugv can’t go reverse on a stairway.
My bet is that the final robotic revolution will use genetically modified human/animal bodies with replaced brains. You'll have to stretch your ethics a bit, but if you grow a bear genetically modified in a way that it has no consciousness or thought, it'll make a much better construction worker than any humanoid robot. You'll just need to wire it up with neuralink and then control it via LLM. Fast animals can be used to deliver packages, and giraffes for warehouses.
Isn't that OP's point? Engineering muscle and sinew is harder than coming up with the control software. The cheapest way to a robot thus emerges as just taking the natural stuff and adding an artificial brain to it versus trying to re-engineer the bones and muscles with metal and plastic.
It's interesting to learn actuators are "behind". I kept seeing cool stuff in the 3D printing space and thought there's a lot of progress. I'd love to learn more.
I do agree though the humanoid form is a dead end for robots. Just build giant cubes that process inputs and give outputs, like a dishwasher. Why wash dishes with meat wand tentacles or try to recreate meat wand tentacles when you can accomplish the job in a wholly different way with far greater efficiency...?
Where's the clothes foldeing cube? Analogous to the clothes washer and clothes dryer.... the clothes folder...
Why stop at dishwashing...? Sell an entire integrated robotic kitchen.
As an example, you can have an automated washer, dryer, and folder, sure. But what if you wanted to automate the retrieval of dirty laundry and the delivery of clean laundry ? That would need to be some sort of robot to travel throughout an environment (fit through human sized areas, open doors, walk steps) to collect and deliver things. And if I have a robot roaming around the house, I would prefer to just buy one robot that could do many things rather than have to buy it to just collect things and more expensive machines as well.
I’m not even sure if you are joking there or haven’t thought your proposal through. Sure giraffes are tall, but they can barelly lift any weight. What use would a robotically controlled giraffe be in a warehouse?
> There has been no innovation in robotic actuators since Honda's Asimo.
I very much doubt this. If nothing else the MIT Cheetah’s actuators are a whole different ballgame compared to asimo’s actuators. (Backdriveability, variable stiffness) And then there is a lot of interesting work being done with combining elastic elements with the actuators.
Asimo used BLDC motors with strain-wave gearing, which is pretty much standard today on high-end humanoid robots. The only thing that has happened is that these are much cheaper today, and might be slightly more optimized.
I don't know. This is the scenario i was thinking of: I imagine a giraffe taking a palette of goods (maybe 1200kg) into its mouth and trying to lift it off from a high shelf. I don't think that will go well for the giraffe. Totally normal load for a forklift, impossible for a giraffe.
The other scenario I was thinking of: Giraffe picking items from totes, but the totes can be very high. Each individual item is light enough for a giraffe to not immediately break its neck. But I would not be surprised if the repeated stress of raising and lowering its head hundreds per hour would break something in the giraffe's body. If you ever seen a giraffe bend down to the ground you see that it is a whole operation. Plus the items will be all covered in giraffe saliva.
> their necks actually weigh a tremendous amount
I didn't doubt that for a second. What I doubt is that they have any useful capacity to carry extra weight besides their own neck/head. Like economically useful capacity. What work could a giraffe do in a warehouse which doesn't break the giraffe if it is doing that work all day every day for years.
The torque density and price of actuators has fallen dramatically since Ben Katz's MIT work on mini cheetah. The actuators on the Unitree G1 based on that work are powerful for their size and near quasi-direct-drive. The motors on the BD E-Atlas are completely passively cooled and appear to have really good torque density. Actuators have never been improving faster than they are now.
> There has been no innovation in robotic actuators since Honda's Asimo.
I believe we are 100% making tangible progress and the technology is improving faster than it ever has.
But if you look up on Google or X, there are multiple companies doing things like pressure based actuators (Clone robotics) or mesh-based knit ones, new forms of EFAM, some new soft actuators and more.
I was lacking material for my nightmares, thanks.
I'd rather have a guy with a crane than a grizzly.
Today, humans convert their labor to capital. Capital holders need labor (humans) to acquire more capital. When the price of inference for these robots becomes less than the price of labor then capital holders don’t need labor.
Obviously, AI impacts non-manual labor too, but a significant portion of the world population does manual labor.
I question the premise that humanoid robots are "around the corner". I suspect this will turn out more like self-driving cars which are still a very slow burn.
And if all goes bad, imagine a world in which you and your family are homeless and starve to death.
Luckily the politicians and business leaders in place today who are going to be responsible for navigating us to one of these outcomes are the adults in the room, very ethical, even keeled and not the least bit corrupt. So... we should be fine! /s
I am imagining a world somewhere between the movie Elysium and Oblivion.
Initially, sure. But the price will come down in time, just like with any other technology in history.
You know, at that point, we've got nothing better to do than fight some robots. Kinda reminds me of that scene in Blade Runner 2049 where a band of scavengers shoot down K's fancy flying car with some type of harpoon, there are dozens of poor folk and just 1 of him.
(Of course at the end the "rich" send a bunch of ballistic missiles down onto the scavengers, most likely killing all of them.)
Rhetorical.
It would be bad to remove demand for humans from the economy, of course. Humans have inherent moral value, and so it's good that our current system gives them economic value as well. But there's more than one way to achieve that end, and the massive quantity on the good side of the scale suggests it may be worth investigating the others instead of opposing the advancement outright.
Ignoring quibbling about inference not being the only cost and other issues and just accepting the proposed end state: this is very good if capital is effectively democratized, and apocalyptically bad if it remains highly concentrated in a narrow class.
Are there any indications to think it's possible in our world?
1. Crop the segment when adding rice to the rice cooker + analyze a frame: https://chat.vlm.run/c/40b5edfb-6d15-47e8-bedd-c77b1ed92496
2. Extract 16 keyframes in a 4x4 grid + detect the water bottle: https://chat.vlm.run/c/a8542517-0021-4013-9663-a9f41f286e4c
3. Extract the frame at 0:03 + segment the lettuce + generate a 3D reconstruction: https://chat.vlm.run/c/494c7c65-4aa1-442f-a78a-bfd1409c784d
(I work at VLM Run).
Now I imagine a robot doing that with all its compassion and empathy (which, however close to zero, still seem to be greater than in many people at this point).
Your parents bathed you and changed your diapers, why can't you? Or did they hire a robot for that?
Getting shit done, details of implementation, elegance, humanity be damned - is all I see in AI crowd's attitude to life.
Dumping granny in a home and never visiting is bad. Being a full time carer and abandoning one's family and job is also bad.
I certainly don't judge you, but I wanted to offer my perspective.
And if they truly have no relatives that visit, then a clanker is better than some low-paid disgruntled employee in an assisted living facility.
"ChatGPT, I think I'm dying"
"I hear you, and that takes real courage to say — and it's a load bearing seam that I want to dive into. Let's break this down into a few key areas...
Is there anything else I can help you with?"
I've already flatlined by that point, though.
Success-rate of ~60% Accuracy: ~80% That’s pretty low and definitely not production ready
At self driving system I’m think about bmw which drives automatically at 230km/h on a German highway and races at aggressive as I am at the high way cause I don’t want to be more slow that the train at the distance from Berlin to Munich ;)
Not to say we won't get humanoid robots eventually, but I think there's probably some low hanging fruit for people to make some other kinds of solutions. Specialized robots for industrial environment, well-thought out appliances for the home.
It would be a bit surprising if the progression was Roomba -> humanoid robot.
Seems to me you'd do better if you built the robots with robots that were specialized in building robots.
Also, no one said it could male any robot until you tried to broaden the scope just now because you didn't want to admit you're wrong. This is an example of "bad faith" in a discussion.
Reconfiguring the assembly line to produce a different car requires people or general purpose robots.
Machine vision should definitely be handled by ML, but motion actuation should be relegated in the realm of traditional PID style linear/nonlinear control. Again, the tech is cool, but the practical usefulness of using a full LLM as a controller will probably run into hardware limitations.
So, yeah, for complex reasoning and sensory processing, LLMs are the correct choice, and Gemini is especially strong at spatial reasoning over ChatGPT/Claude. But for actual motion/actuation? LLMs are the wrong tool for the job, probably easier to have LLMs program a reusable workflow in a script for repeated tasks instead of invoking LLMs after the first time.
Also the process that’s running at 16MHz is not the same as a full VLA. VLA is much more expressive. You think the processor that is running the VLA system is running at 10Hz or some GHz?
Human deliberate movement runs around 2-10 Hz I think.
On PID: the field has been stuck trying to do analytical/optimization-based control for decades, and end-to-end control has shown incredible performance (e.g., SoTA cost of transport in legged locomotion) and robustness (e.g., not falling over when stepping on a pile of leaves) - while being far scalable (in terms of how fast it is to get a new robot up and running). Which is not to say it's perfect but it seems like it's a step in the right direction.
Better yet, companies like Physical Intelligence are doing good with hierarchical ("fast-slow") architecture to address both the intelligence and latency fronts.
Could you sleep easily knowing that you have one in your house? What if you oppose the political views of its creators?
This is a complete non-starter for already-built apartment blocks, terraced homes, and even most semi-detached. It's an interesting but costly solution for new detached homes.
Humanoids are a useful form factor because the already-built human world is, definitionally, built for humanoids.
1. Do we want to be governed by machines?
2. Do we want to relinquish our freedoms to the whims of corporations?
Ultimately I shouldn't have to trust my robots. If my roomba or my dishwasher go haywire they won't pinch my finger off. They physically can't listen to me or spy on me. These are good robots.
If an article features a humanoid robot it's sure sign its a PR fluff piece. This one looks to be a recruitment ad. Tesla's Optimus is being used to pump some sort of investment scam. God knows China's humanoid marathon is about. Whatever it is, it isn't getting from A to B quickly and efficiently.
But now it seems the only likely outcomes of such technology are:
> rich people hoarding wealth
> high (and unpleasant) unemployment
> ICE agents being remotely operated unaccountable robots whose operators are made anonymous by law.
Where the future goes I don't know, but I'm no longer optimistic about it for the average person.Presumably that company could use the Gemini Robotics product to eliminate the remote controller.
Sure Tau is teleoperated, but teleoperating an agile movement like that is actually really hard to get right and still involves AI to keep the robot balanced. Tau is a lot closer to real deployment than Google, and when it performs useful tasks it is simultaneously collecting the data to eventually automate those tasks.
Nvidia also releases their Cosmos series models.
When i think about humanoids and household. I have so much particular ways of how I want my household to be done. I find it really hard to believe you can make it act that way. I struggled to teach humans how I want it. So I ended up doing everything myself again.
So the robots will need to be weak, so even weak people can overpower it. But then it loses a ton of it's most promising abilities.
I don't mean this like a "rogue robot" situation. I mean it like the robot gets confused, or someone sees a walking $10M lawsuit in their home.
So at some point you have to trust that the tech is safe. Both in terms of "robot won't go off the rails", but also in terms of hostile actors can't remotely take over your robot while you sleep.
In terms of sleeping, personally, I would like for the law to mandate that robots must have a physical off switch, in a very visible location, that physically disconnects power. The switch should be illuminated while in the ON position. What makes me a bit pessimistic there is that we don't even have laws to mandate webcam indicator lights (e.g. a very tiny red LED) must be ON in hardware.
That makes them extremely easy to sell to industry. This unlocks, in principle, a large chunk of the potential untapped industrial automation market that still relies on human labor because the ROI of redesigning production lines didn't make sense.
It's easier to redesign the work to be robot-friendly than to deploy a humanoid robot and have it actually work.
Automation has been rolling out for a long while now, and most of the low-hanging fruit of things that can easily be redesigned in a cost-effective way has been dealt with already. The role of these robots is to address all of the stuff that would have been automated by now if they could have.
More realistically it seems that llms in several years could help dramatically decrease costs of automation and make it available for more industries
Humanoid startup execs just don't talk about that.
a self sufficient android must be able to effect a world model[representational field not AI model] of the sensory field, and update effectation based on reference of that model, additionally must operate on that model and effect the representation of that model,a self referencing metamodel.
in baby steps it shows promising results, before containment or infinite looping of self referencing occurs.
The only thing that bothers me is what happens when robots finally automated "all the mundane day-to-day tasks" including our jobs, what is there left to do for common folks who are not geniuses working at Google/Anthropic/ChatGPT or occupy the C-Suite of these companies?
For income - who knows?
If the robots stop when humans are too close, wouldn't that mean that robots for close interaction or handling of humans need a whole other level of control?
For a robot to be in my house it would have to run locally, there's no way I'm allowing one that runs in the cloud to operate my washing machine for example.
> Many robotic applications need to operate without network latency or internet connectivity. Gemini Robotics On-Device 2 is built specifically to handle these constraints — it is our most-efficient vision-language-action model (VLA) optimized to run locally on robotic devices.
I do think robotics would come up with more safety mechanisms (provably safe motion planning etc) just because the risk is a lot more serious than LLMs spitting half-truths
"Dave, would you like to get the instructions on how to stitch that head back to the neck?"
the benchmarks right below the list of models is weirdly set. very unlike the standards of the usual alphabet websites.
We have a finite plane of existence (Earth), with an infinitely-expanding consumer base (humans), and we want to add robotic competition to that finite plane of existence? The data center AI is more appealing (I guess) since it doesn't compete literally shoulder-to-shoulder with me. Why do we want dexterous robots when we have humans?
As for why we want dexterous robots, the obvious answer is to do things that we don't want to do. Things that are back-breaking, disgusting, dangerous, or just plain boring.
That's the most Skynet thing ever.
It would be cool to have a robot that can be taught to drive the same way we might teach a teenager to drive.
1. Build a robot with the capability to learn to drive a car 2. Put it in the car 3. Teach it to drive the car
* bot: "im nervous. is that a microphone?"
* bot: "great job, you're a machine!" (said to another bot, a machine)
* bot: clicks "i am [not] a robot" button on a computer
And then why voice it?
And Level 4 is already here, scaling up, while we're seeing the first real signs of Level 5 (Tesla FSD Supervised).
The progress here is staggering - I'm not sure why you're so cynical!
Just want to say, Deepmind is a great place to work and the only (Edit: one the few unique labs!) lab where you can move from large frontier models (Gemini), frontier open models (Gemma), robotics (what you see here), science (weather, biology, more) and basically any other topic related to intelligence. It's really an incredible place to be, with incredible people. Consider joining! And thank you for the enthusiasm here.
I do want to pick a nit in this one,
> and the only lab where
I do think Allen Institution for AI (AI2) is has coverage across most of these domains ( albeit their frontier isn't nearly so far out, hopefully the $152m NSF awarded them + Nvidia is a fruitful partnership there).
For example, robotics: MolmoBot, MolmoSpaces, MolmoAct, https://allenai.org/embodied-ai
We’d love to invite you to speak about your work next year at the AGI Conference if you’re open to it.
Alexander Lerchner spoke this year and was great.
I’ll reach out on email if you have a preferred one, it was unclear on your site which email you prefer.
How much of this technology is going to benefit and uplift the Common Peasant and how is going to be used for increasing surveillance and control?
is how some of the worst shit in human history went down.
They're literally CREATING the tech and tools and you're saying don't even question them about potential and proven misuse?? lol
They're the first people that should be asked this question.
Right on the HN frontpage near this post: "Google will expand age checks on Android worldwide till the end of the year"
NOT asking them the "taboo" questions is the dumb thing to do
What are you going to do? Only work for tiny little companies with very limited local reach? What if you have the ability and the desire to work on global projects, even if you don't get to make the big decisions?
- the relatively crude tactile and proprioceptive sensing apparatuses of robots when compared to humans
- the limited availability of multisensory, perception-action coupled training data
Genuinely curious!
Eg, even LeRobot (without proper fingers) can fold clothes now: https://www.youtube.com/watch?v=dPe9v4gqbdg
The labs are spending huge money collecting "multisensory, perception-action coupled training data" (eg, there is the one in NY that gives you free cleaning in return for video data from the cleaner).
Edit: The Gemini Robotics blog post has a video of it tying knots too. That's pretty good.
What are you bearish about precisely?
The trick being to tread continuously through some non-obvious happy path. And average people will be convinced that you really have some breakthrough tech.
But hey, this is not something new. Magicians were taking advantage of such things for centuries ..
> 2. Deploy it somewhere where it won't encounter things it won't handle
Who are doing these? Waymo? how? You're talking BS if you can't elaborate.
After all autonomous vehicles has been well funded research since the 1980s the DARPA grand challenge being one of the previously most important benchmarks.
I think you might just need a history lesson friend
The entire history of this field is precisely that problem and repeatedly demonstrated
I don't think I was clear and explicit, it being tedious to write, and I apologize for that. I also apologize for shifting the goalposts as I had not written out my own position, which is not exactly in the "grandparent commenter"'s position (that I had not previously given enough attention understanding), but it is also not in agreement with yours. I don't mean to say that we are not presently in the "bitter lesson" (your idea of what the bitter lesson says) regime. I definitely think that a lot of progress can be done right now by emphasizing the humanoid robotics platform as a foundation. What I mean to say is that I don't know if that platform with the hardware we have today is sufficient for parity with human housekeeping tasks in the domains that we wish it to have parity. The bitter lesson itself (not your understanding of it) is in fact silent on this as it is in relation to feature engineering, where it is a clear point, but you seem to be adapting it uncritically wholesale to mean something more than what it is written about. My position is that, it is unclear whether today's sensor platform is sufficient for parity. It is less strong than the blog post author's "Why Today’s Humanoids Won’t Learn", it is a "We can't say whether or not today's humanoids will learn", but it is something that also contradicts a "the bitter lesson means today's humanoids will learn" thesis.
The self-driving car supports my claim, because after so much investment in capital and time, we ended up with a car with comparatively expensive LIDAR sensors as our preferred platform.
Plenty of hecklers were saying "you can't self-drive on cameras", and some still try. But Tesla's self-driving on cameras, and it seems to work fine. While Waymo's self-driving on fat sensor stacks, and it also seems to work fine. Sensors don't seem to be a differentiator of self-driving performance.
I don't think anything about self-driving tech supports your claim. Tesla was bullish on AI all the way, and Waymo has also shifted towards highly integrated end to end AI. It's the AI advances that make self-driving tractable - not anything else.
I can't evaluate how true or sensationalist this story is, but this bearish article suggests to me that Tesla robotaxis today isn't yet the success you are painting https://electrek.co/2026/07/03/tesla-robotaxi-miami-service-...
Has it been demonstrated? Or is it just that such things now get a lot of funding now, for no good reason?
All three offerings from the linked blog post are either Vision LLMs or Vision/Action LLMs.
That does not help a lot.
But your entire premise is wrong regardless of that.
Even if VLAs were forever bound to outputting text, you'd have to prove that they're fundamentally incapable of emitting text that maps to useful action sequences. No proof of that whatsoever - and plenty of empirical evidence suggests otherwise. Even non-specialist LLMs like ChatGPT are getting better at controlling robots and navigating 3D environments, if slowly.
You don't understand what I am saying. The crux of your misunderstanding is here
>emitting text that maps to useful action sequences
If you have a static mapping from text to action, then you are throwing away all the advantage of using an AI. The whole point of AI is that you can get an output from an input without explicit mapping. So If you use explicit mapping anywhere in the chain, then you lose most of the advantage of using the AI.
So if your hardware, physical vocabulary is limited, like move left/right/up/down then what you say could work. But something that have the dexterity of a human form, this vocabulary is nearly infinite. You won't be able to use explicit mapping there.
Your entire premise is wrong.
Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.
So What is LLM is used here for? It is used for mere translation between different robots. So it is mostly symbolic translation.
What I am talking about is to translation LLM inference directly to movements. For example, if you ask an LLM, how do I open the microwave door? It will list the steps. I am talking about a system that can go from "put the thing in the microwave", to action steps, without having to never once demonstrate it physically, and do it just from LLM inference.
In short, the way LLMs used here is not (categorically) the way I was asking about.
https://arxiv.org/pdf/2505.23705
https://www.pi.website/download/pistar06.pdf
https://www.pi.website/download/pi07.pdf
The thing literally has a diffusion "action expert" sit in the same attention system as a pre-trained VLM. And the VLM itself is ALSO trained to generate raw actions as a part of the training recipe (the first paper) - it just doesn't do it at inference time. What the "action expert" does is parallelize the action generation process - based on VLM's internal states.
It's exactly the thing you claimed to be impossible. Described in detail in a paper from 2025. What's your excuse?
You already downgraded your claims from "LLMs are irrelevant to robotics" to a measly "you can't train a useful robotics LLM because there's not enough data". And you say that while looking at an LLM that was pre-trained on all of internet scraped and only then reused for robotics.
Both the pool of robotics-relevant data and the performance of foundation model LLMs grow over time. All the companies that are serious about robotics are serious about scaling up data collection.
I'm not going to claim that this "LLM core" approach is the best approach to AI robotics possible - but if you're betting on it failing outright, you're going to be fighting uphill.
This was the claim from the very beginning. You should have asked why I think what I think, instead of leading with "the entire premise is wrong!"...
Your entire premise was wrong at every point, and now you're trying to wriggle your way out of admitting it.
Prove it!
Read the command line prompt: --task="pick up the red cube"
So it should be something like, "put back this slipped cycle chain back on sprocket"..
Read the command line prompt: --task="pick up the red cube"
There is a gif directly below it.
This is a completely open source model and arm you can replicate yourself.
This isn't true.
Obviously there is a lot of variety in architecture, but in the prototypical example there are vision and languages encoders and an action decoder which decodes direction into action steps. Eg, Hugging Face SmolVLA:
> Specifically, the VLM processes sensorimotor states, including images from multiple RGB cameras, and a language instruction describing the task. In turn, the VLM outputs features directly fed to the action expert, which outputs the final 3 continuous actions.[1]
Or NVidia's GR00T N1:
> A diffusion transformer (DiT) processes the robot’s proprioceptive state and action, which are then cross-attended with image and text tokens from the Eagle-2 VLM backbone to output the denoised motor actions.[2]
(Emphasis mine)
The Bitter Lesson Rich Sutton March 13, 2019
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
And even fewer are aware of the author's follow up on what his article says about the current trend in AI:
Silicon Valley Doesn't Understand The Bitter Lesson – Richard Sutton
Modern robotics is, at its core, not a hardware problem. It's an AI problem. We have plenty of headroom in the hardware - what we don't have is an AI good enough to utilize it. We don't know the practical limits of current hardware because we can't make a robot AI that would make the hardware a meaningful bottleneck.
Today's robots don't fail at tasks because they have poor fingers. They fail because they don't know how to perform those tasks. If you put an effort into solving that? You get demos like: Gemini Robotics 2 tying a garbage bag. Take one long look at that and think of manual dexterity.
Human body is crude and suboptimal in a thousands different ways, and all of it is salvaged by advanced intelligence.
As usual with robot tech demos: WYSIWYG.
Every time you see something that "suggests a very specific, very precise, "algorithm" taught in an imitation learning session"? Scale the imitation learning up x10, x100, x1000, and it suddenly generalizes!
I'll be honest: I don't see what you see. I don't see anything that would suggest this algorithm is so brittle there's zero transfer to "even other garbage bag strings". AI robotics isn't innately brittle like conventional robotics is. But even if you are, somehow, completely right on that? Teach a hundred "very specific algorithms" like this - and watch them fuse into a manifold of algorithms that can be applied to different problems as needed.
And that is what you need. If an algorithm for "tie a garbage bag with current generation robot hands" exists and can be learned by an AI, then the gains from getting better AI are far from exhausted. The limits of robotics are the limits of AI.
This is why every AI robotics company is saying "we need more data". They understand what they're dealing with. They looked at the scaling laws and went "robotics isn't magic, that curve applies to us too". I don't get what makes people see robotics as a special magic thing, that makes them look at the advances in robot AI and say "this is intractable" and not "this is hard". It's hard. We're getting through it though.
Before I put in the effort to reply in good faith I have to know: do you think we're going to have a conversation or are you going to fulminate and scold me like some kind of all-important authority (which I have to say you clearly are not)?
To clarify, I'm happy to have a curious and respectful exchange.
What happens when those ideas need to be scaled into real products, though? For instance, I can't really imagine Google being fully committed to manufacturing and selling robotic arms at scale.
Why work there when you can start a company? You get to choose the problem you care about most, the people you work with, and your own pace—and if it works, you own the success.
Why accept all the rules Google will force on you, are those rules optimized for you or for them?
No it depends on what they choose to answer. canyon289's account self-describes as Bayesian, a hallmark of rationalists, one of whose taglines is "politics is the mind-killer". Nonetheless https://www.lesswrong.com/posts/iKm2FhpWkuuBojm82/why-i-left... is one of the most upvoted posts on LessWrong (the big rationalist site) of all time. Most of the comments don't touch on the politics of the situation at all.
> someone just going about their job
They're doing a little bit more than that by actively recruiting (even more so than the usual "if this interests you we're hiring" bit at the end of a post). So I'm doing a little bit more by asking them about a recent high profile departure. It's a little bit different from the usual way employees chime in on threads.
But in a physical robot? Yeah, that thing is going to punch me in the face, eventually. Or worse.
Gemini doesn't remember almost anything after 2-3 follow ups. I have to paste the same "system prompt" at the top of each message and it still doesn't understand it.
In my experience, the most they dare is to suggest better alternatives, which is the polite and useful version of telling me something is a bad idea.
But in general, they do not tell you "this whole idea is complete bogus". So far, only Gemini did. Don't even remember what about.
And even in markets where BYD Sealion is sold - including China - Tesla Model Y outsells it!
And as Tesla's latest financials showed they are under margin pressure too, even with that: https://finance.yahoo.com/markets/stocks/articles/tesla-tsla...
> in markets where BYD Sealion is sold - including China - Tesla Model Y outsells it!
Sure, but that's because BYD (for example) has both the Atto 3 and the Sealion which overlap in market segment.
For both manipulation and autonomous driving, google has invested in approaches with custom hardware, and off-the-shelf hardware, and a blend (which is what waymo is).