Figure 01 robot demos its OpenAI integration
twitter.com
twitter.com
It's really interesting to see it integrated with a robot that can interact with the world though. I think that what's really holding back the current crop of Gen AI is inference cost and speed. When we figure out to get thousands of token per seconds for cheap, I think we will be able to bruteforce many hard problems and actually start seeing amazing applications like this one (but in production rather than a cool demo).
Can you imagine seeing this ten years ago? Moving so far in such a short time frame would have been unbelievable.
This gave me the "I'm actually living in the future" vibes that I always imagined I'd get from flying cars.
It's very easy to forget the single line hallucinations of GPT3 not too long ago.
Look into providers like Groq that are running real-time inference with Transformer models.
It's only a matter of models improving then the hardware coming to market that this will be a major leap in robotic capabilities.
Tools can be useful for many things, but tend to be inflexible compared to neural nets.
The way I see it, the main benefit of such tools come when the part covered by the tool is very repetitive and well defined. In such cases, algorithmic solutions may offer better precision and lower power consumption than neural nets.
Anyway, once most of these algorithms are actually WRITTEN by AGI's, the difference may cease to matter.
It's unclear how general these movement policies are though. The way that guy is standing perfectly still makes me think that it would fail if everything wasn't set up just so. I'd like to see demos with more variation.
I don't want to be too negative here though. I think it's a great demo and I can't wait to see more.
Yes. Compare "Put That There" (1979).[1]
> The way that guy is standing perfectly still makes me think that it would fail if everything wasn't set up just so. I'd like to see demos with more variation.
Yes. Unstructured manipulation is hard. Structured robotic manipulation is pretty standard. Picking isolated objects is a solved problem. Here's a robot recycling sorting system, "Max-AI".[2] That's been in use for years. San Francisco recycling uses those robots. So do many other cities.
(That's from "Bulk Handling Systems", a company which does exactly what their name says. Recycling and trash come in bulk, and their machines handle it. Shakers, magnets, screens, and vision-based air sorters do 95% of the sorting. The robots only handle the hard cases. This is the no-bullshit end of AI.)
I'd love to be wrong but I expect that if they had that much flexibility in their controllers they would have demonstrated it.
Covariant's RFM-1 (announced Monday) is an interesting approach to generality. https://www.youtube.com/watch?v=1Go6HEC-bYU
At 0:59 it also shifts the basket out of the way as a separate action, after its already put it down.
Here's a demo from a startup in this space. Still very early. https://deck.sindarin.tech/
Also my radio-trained voice is so generic a caller every week-ish assumes I am a bot, so I’m pretty sure the problem isn’t me enunciation or accent.
> [INST] Spell out the following word letter by letter: margorczynski [/INST] m - a - r - g - o - r - c - z - y - n - s - k - i
> So, the word "margorczynski" spelled out letter by letter is: m-a-r-g-o-r-c-z-y-n-s-k-i.
The text between `[INST]` and `[/INST]` is the input. The text after `[/INST]` is the output.
https://twitter.com/karpathy/status/1657949234535211009
I'm not arguing that you can't use single chars just that many of the issues parent discussed are caused by this.
Aye.
I was surprised this morning when it decided I had was talking about a "Mark of chain". 1/3rd of the time it hears "bedroom 100%" as "bedroom off".
When cooking dinner today, I asked for a "ten minute timer", it responded "for how long?" then confirmed my "ten minute minute timer".
Still better than Alexa, which kept telling us it couldn't find «kitchen» on Spotify even though we didn't even have Spotify.
And way better than the voice control on Mac OS Classic; back in the late 90s/early 00s, it interpreted 75% of my attempts to use it as "tell me a joke" (it wasn't even a good joke), and ignored 20%.
This demo is pretty bad compared to what we currently have in development.
We’ve been in code freeze in prod for over two months to get our substantially improved engine finished.
It’ll be out in a few weeks, and it’ll blow this version away in every way that matters.
Thanks for checking us out!
The vast majority of America is within 10ms of a data center. That's nothing.
The current challenge for most interaction is ASR -> prompt processing latency. This will be improved with multimodal models on specialized hardware like Groq.
and another 0.6s or so to get first voice chunks from PlayHT
measuring STT latency is harder, I need to implement a local VAD model first to properly measure it, but I think it's on the order of 0.5s
So this has nothing to do with Groq, really. ChatGPT is just slow (too slow for realtime voice communication).
Why do we add junk words while we think? I think it's probably because we're social animals, we want to hold that person's attention as we think as periods of silence are likely to make them become disengaged. But who knows really.
Can you call this an AI wrapper company? Kinda! The medium is a little different than an app, of course.
Lots of amazing applications of AI even if frontier AI development froze today.
yeah this was super impressive. If this is at the point where you can put an arbitrary object in front of it and ask it to move it somewhere, that's going to be huge for industrial automation type stuff I'd imagine.
I do wonder how much of that demo was pre-baked/trained though. Could they repeat the same thing with a banana? What if the table was more cluttered? What if there were two people in the frame?
Boston Dynamics has been demoing the pre-baked dance routine for 2+ decades at this point. Really hoping we can evolve past it.
"Take 488... action!"
Not sure anything here is state of the art, but that doesn’t make it easy.
It would be absolutely amazing if they really are at that level of manipulation in general, and it would put them vastly beyond what anyone has been able to do date. However robotics demos have a great history of being a mix of slight of hand (partial/full tele-op), heavily cherry picked, or tuned to a extremely specific example.
Because it's such a leap being implied by this video, it's reasonable to want significantly more evidence before believing they can do this type of interaction and manipulation in a general way. But even if it is heavily leaning say on imitation learning for this exact scenario, there are tons of potential applications for this level of capability.
This stuff is not infeasible, these robots can be built for the price of a car. There is still a ways to go from this to androids, but it is legitimately technologically feasible now.
It sounds so human, a person would also stutter at an introspective question like this. I wonder if their text to speech was trained on human data and produces these artifacts of human speech, or if it is intentional.
The naturalness of the speech is extremely good, though.
I noticed the stutter too, interesting to see that is what TTS just does now, and not a sign of a human and a sound filter.
Truly flabbergasted.
HN you're not supposed to be this gullible...
But this guy is a professional advice giver, so to be expected?
Wouldn’t surprise me if he outsourced his tweeting.
You can sign up and use their Voice Lab for free or maybe a few bucks and experiment with the slider for Stability and the other setting.
In my opinion, turning Stability down just a little bit to demo extremely realistic speech is a no-brainer. They could have turned it up and made it ultra-smooth, but that makes no sense. Why make your robot demo less realistic deliberately?
Figure Raises $675M at $2.6B Valuation;Signs Collaboration Agreement with OpenAI
I'm sure it's a bit cherry-picked and they chose things it is good at. But it is already showing some useful stuff.
(All usual caveats of stochastic parrots and non reasoning apply)