Conversation with Zuckerberg, this time we talked as photorealistic avatars
twitter.com
twitter.com
This is actually impressive, and while not perfect, is close enough that they'll probably reach a version that is viable very soon.
Typically you do one MVP version and as soon as you've knocked that out, start working on a proper version, integrating feedback from the MVP as it comes in. I'm guessing they always wanted to have photorealism, but realized it'd take a while to get in place so they launched a quick version then got back to work.
This is not a phone app. Physical demos that propose something that a user has no point of comparison with (new paradigm) absolutely need to be on either end of the maturity scale.
- Either very raw barebone MPVs that should be only for lab/research aware people who won't cling to details and are able to see the 'principle' of what it could be.
- Or absolutely imaculate picture perfect marketing demos that you can show to non-technical people
Anything in between is a recipe for getting shot down, getting no buy-in, and receiving no insightful feedback.
Really? Definitely looks like something a simple phone should be able to process... Are they not wearing phones on their heads in the demo?!
This "Publish > Feedback > Iterate > Repeat" loop doesn't just apply to smartphone apps, you can basically do this with anything in life, from hardware to mixology.
I guess you could say the same thing about LLMs, since the initial models were terrible and just basically demos, they shouldn't actually have published those at all?
I again disagree.
"Publish > Feedback > Iterate > Repeat" loop works when users can build on it from a previous reality that they can compare against.
Physical demos that propose something that a user has no point of comparison with (new paradigm) you can either show a sketch/concept of what you want to do (to a tecnhical audience). Or build a semi-realistic demo to a non technical audience.
Anything in between will probably be met with confusion or people focusing on details that are not the concept that you are showcasing.
I don’t agree. I dislike Fb to the extent that I already deleted my account quite a while ago. And I thought the Meta attempt at a meta verse with cartoon characters looked ridiculous.
But the video in the OP link, with the photo real avatars. This is a whole other level. It’s super cool.
The fact that they made a ridiculous cartoon Metaverse first does not detract from the coolness of this photorealistic thing.
Still, I’m not gonna use any Metaverse that Meta will make. Photoreal or otherwise. But that is because I dislike Meta as a company.
The point is that what they have made here is some top notch stuff.
I look forward to seeing Apple compete with this. I hope that with an iPhone and an Apple VR headset I can get to experience something similar in a few years.
I agree with you on the coolness of this demo. I disagree that the cartoon demo does not detract.
I suspect a large majority of users and advertisers assumed Zuckerberg went off-piste completely. This interview with Lex is much more aligned with what should have been their first public demo of a Metaverse 'vision'.
edit: looks like Epic is working on support for VisionOS in UE5 https://github.com/EpicGames/UnrealEngine/tree/ue5-main/Engi...
Haven't we seen this type of thinking to be proven a fallacy enough in the last decade to at least be skeptical of solving that last 1% until we actually see it?
Rendering two fairly realistic models on the screen at once isn't a big challenge these days, if they're the only thing being rendered.
I doubt it's a performance issue: the 3D model of a rectilinear room has fewer vertices than used to represent a single earlobe. It most likely is just an unimplemented feature; they also could have "cheated" by chroma-keying a 2D-texture background or rotoscoped the avatars onto a another 3D room scene, but what's the point when its just talking heads?
Games often only have the highest detail models present in cutscenes, where the rest of the game can be limited, and the camera can be controlled to limit how many characters are visible.
Yes - they do, but you don't have to get fancy with the environment for 2 static avatars locked in place. What are the odds that rendering a room was dropped because it is allegedly GPU-intensive vs it being a feature that is not fully-baked for a public debut yet?
Totally agree that it’s a stylistic and presentational decision and not a computational limitation.
That demo was going to happen one day or another as the tech progresses, and it'll only get better from there.
But my point was more really that even if they paid him 10 million, the PR would have been worth it.
>That demo was going to happen one day or another as the tech progresses
Yes, all amazing technology is going to happen eventually — no one really cares about that. They care about seeing it happen in the present, right in front of their eyes. The metaverse was mocked profusely in 2023, and we have gone through multiple cycles of “This is going to be the year of VR” in past decades. Tech stocks have been flat the last couple of months and Meta is a public company. No one on the outside would have batted an eye if they quietly shut this whole thing down after being so relentlessly mocked for the past year. So saying Zuckerberg doing this demo was inevitable after the fact is a bit flippant I think.And if facebook did a demo with a bunch of celebrities and put it on their website I’m not sure it would have the same effect as this long interview with Fridman. It wouldn’t have the same effect on me that’s for sure. And I’m pretty sure they would have paid those celebrities.
Of course anyone in the real world realizes how little credibility these "interviews" have.
There's a time and a place for that.
It does not mean they will make a killer product. But behind closed doors they have some amazing tech.
No real journalism involved, just talks about peace love and prosperity to some of the shadiest, shittiest businessman and politicians of our time.
I agree with the earlier part but disagree with this part of the comment. If the interviewee is someone particularly interesting, I learn a lot more about them when the interviewer just throws softball questions to help the interviewee more or less free-associate. Most of the time when an interviewer tries to "challenge" or "ask hard questions" it comes across as cringeworthy and a waste of time that could have been better spent letting the interviewee speak their mind.
I actually like his interviews, because he typically speaks less than his interviewees and doesn't guide the viewer to a certain opinion. Even his interview with Kanye West was like this - it was pretty insightful into Kanye's state of mind and didn't need any commentary.
I find he drones on way too much, and way too many first-person pronouns. "I-me-me-me-I-I-I" - it's pretty narcissistic when the guest is politely sitting there across.
Surprised you also brought up the Kanye interview, that's a great example of Lex getting reprimanded by the guest. Lex couldn't let go at a particular moment and kept injecting his opinion that Kanye had to slap down.
I’m guessing if you’re a famous scientist you probably don’t care/mind being interviewed by most poscasters, but Lex is the next-level thing to do
Six years ago I went to an arcade called the rec room in Toronto where they have a ghostbusters AR game I got to experience with friends[2]. To this day I still have vivide memories of the experience. When it was all done, we all sat around for 15 minutes in complete silence, just recovering from the experience. I understand the skepticism about the metaverse, but after that experience there's no doubt in my mind it's the future -- you really need to experience it to get a sense of what a more polished version of this feels like.
[1] https://blogs.nvidia.com/blog/2020/10/05/gan-video-conferenc... [2] https://www.youtube.com/watch?v=ar9YwEv2ACk
IMO the point is not if you will ever think VR is promising or not, it’s about when you will have your first mind blowing moment playing VR
The interview gets quite interesting nearer the end where Zuckerberg talks about his vision for the future, like training AI-backed avatars to simulate real people. Which apparently is not far off from reality - they joke about Fridman doing a podcast where he interviews an AI version of himself in the near future. Reminds me a bit of that Black Mirror episode where Miley Cyrus' character has AI versions of herself contained, or trapped, within consumer electronics.
It also occurs to me that this could be a useful technology for people with social anxiety who have to endure uncontrollable blushing when conversing in real life, or people with awkward facial tics. These could be filtered out or not even recorded before being encoded for transmission, while still giving a realistic appearance, making for a more comfortable presentation of the real-virtual self.
They should open up the Codec avatar creation - lots of people have the hardware and the time/expertise to create them
Each scan captures dozens of terabytes of data and lasts for hours, at least for the current high resolution avatars
You can see from the normal map in the video it's pretty detailed but at least for a single face capture you can do this at home. I'm not sure what secret sauce they have for capturing multiple facial expressions and some ML magic how to morph/animate between those.
Btw you can get good results with things like https://www.unrealengine.com/en-US/metahuman too
This vision of the future sounds like a nightmare to me.
Basically for each of the identified poses you have a key and you grab a pose of it (photogrammetry here) and then you would capture performance and identify somehow which combination of keys and weights would be translated to you model. Sounds easy but it ain't.
The technological showcase is really cool though. Right now people are paying dozens of thousands of dollars for fully rigged vtuber avatars and 3D virtual chat models, but in a few years we'll probably have AI tooling that allows you to do it yourself.
The discussion of non-human avatars made me wonder what kind of fantasy avatar Mark and Lex would use. If profile pictures are any kind of indication, I suspect a large number of tech users will opt to be cute anime girls. The days of catgirl-ification grow ever closer.
They discuss an idea of having celebrities and famous people train AI models so fans can interact with them. That seems so dangerous, arguably pushing parasocial relationships to a completely new level. It feels like it fundamentally hacks the human brain. Are we going to reach a point where most interactions are mediated through various layers of AI models? Maybe I'm being too much of a pessimist...
Now someone do a PGP version of the Codec Avatar, where facebook included can't access the raw data, but streaming is still possible. Otherwise we get the Meta version of Worldcoin and merrily continue on the "we're completely fucked" branch of this multiverse.
Maybe it depends on the person(s) involved? ;)
Bell Labs is a famous example where it seems to have played a role (together with the people working there of course, and other variables).
> But just as important was the culture of collaboration that the company fostered. The leaders of Bell Labs understood that physical proximity could spark innovation, and they designed its facilities to bring experts together in both deliberate and unexpected ways.
> At Bell Labs’ headquarters complex in suburban Murray Hill, New Jersey, all of the laboratory spaces connected to a single, vast corridor, longer than two football fields. Great minds were bound to cross paths there, leading inevitably to spontaneous and meaningful interactions. As author Jon Gertner writes in The Idea Factory: Bell Labs and the Great Age of American Innovation, “a physicist on his way to lunch in the cafeteria was like a magnet rolling past iron filings.” Throughout the labs, employees were instructed to work with their doors open, the better to promote the free flow of ideas.
Easy to gloss over and ignore when they don't support a decision that's already been made.
Yeah I'll pass on that, thanks
I'm not saying all companies should work in a office, I'm just sharing my viewpoint from someone who prefers in office compared to remotely.
RTO makes sense if everyone in the company is working in the same building and sharing spaces for the same reason.
It's exactly the same. You can have the avatars doing anything you want, there is no difference.
Want them to do a pass by every desk in the office when they go to the kitchen, or bathroom? Possible. Want to turn that option off to get down to work? Possible.
It's cool.
Willingly, basically none, or as few as I can realistically manage. I don't have a smartphone other than a burner one with GrapheneOS for my bank app, run linux on everything else and my work MacBook sits idly in the corner somewhere with 0 battery in it despite protesting from my manager.
It would be like Ford employees being largely unable to drive. Which I hope is not the case...
Fucking Zoom declared mandatory RTO, which says a lot about what they sell the rest of us on.
And Ford...heh. Despite incentives, their own employees refused to buy their cars to such an extent that competitor vehicles were banned or relegated to remote parking lots.
That's actually a good thing! It means they took it seriously, and their employees were incentivized to fix the problems so that they actually wanted to drive the company's cars.
If the only way an employer can tell if their employees are working is by forcing people back to the office, the company has bigger problems that employees not working.
I’m sure some people are Apple are freaking out. Meta has a 10X more affordable product on the market.
They are all in on spatial computing and AI.
The future is exciting indeed.
I’m reading Snowcrash again and just been blown by how fiction is turning into reality.
Raven: “You wanna buy some Snow Crash, man?”
Hiro: “Snow Crash?”
Raven: “It’s the most expensive drug there is.”
I've talked SO MUCH trash
I wonder if it could be useful technology in helping face transplant patients get used to their new features, in advance of the operation. A virtual mirror, with a reflection of the future.
If you were to chat with someone daily for a year, you'd never see their hair grow (or get cut) for example but then I suppose we never actively think of these things, just notice when they change.
It's just a matter of time before this stuff is incorporated into what we're seeing here.
[1] https://screenrant.com/rdr2-hair-growth-length-time-tonic-sp...
[2] https://www.watchmojo.com/articles/10-most-realistic-feature...
When they achieve that, it will be quick and simple to update your scan.
The people that keep an old scan at that point are the same people that also keep an old photo as their profile picture anyways, and it’s often not due to the technology.
Some are too lazy to update profile pictures no matter how simple it is. Some people don’t know how to do it no matter how simple, but that’s more rare. And some people purposely choose to keep an old photo because they want to be seen the way that they once were rather than the way that they are now.
Me for one, I will keep my 3d scan up to date every now and then. Say, every few months or so. Just like how I currently update my profile picture every few months on platforms that I actively use. (For example, I’ve been at my current company for about a year now, and in that time I’ve changed my Slack and company GitLab profile pics one time – from the picture I chose when I joined, to a new up to date photo).
Anyway. I am sure that if people stick to old scans, the Metaverse companies will eventually counter that by virtually “aging” the scanned models that people use. So that even if you don’t change your scanned model, the Metaverses will add grey hairs, wrinkles, etc to it over time.
They talked about this in the recording, calling out (not/)shaving and weight fluctuations may or may not be reflected.
My own thoughts: we partially already there, considering people hardly update their static avatars/profile/professional headshot pictures on a weekly basis . We're inching towards the "Residual Self-image" of the Matrix universe
I'm imagining that with actual built environment it will be so much nicer as well. This reminds me of when a friend and I visited VR spaces (in VR chat I think?) felt like I was visiting a minecraft universe all over again with a friend.
Pupil diameter in most humans is affected by autonomic arousal, e.g. in conversation it provides an often unconscious signal to the listener/observer. I didn't detect any dilation or constriction of the pupils in either head image; and for me it introduced some uncanny valley-ness.
The contrast between Mark's lighter iris color and both the blackness and relative smallness of his pupils drew my attention repeatedly. The middle image of the sampled video shows some contrast between iris and pupil but that might have been too noisy for their use. Anyway, I'd be curious, what they tried here. It seems they're rendering the pupil, I wonder if they'd tried playing with the diameter as a fixed proportion of the iris diameter, or whether they tested edge blurring for lighter-colored eyed individuals to reduce contrast.
I'd be curious to learn, but suspect that sending the "wrong" eye dilation information may be worse (e.g. sending a "beady" eyed signal triggering unconscious emotional responses) than just sending a static pupil size too.
Still a very impressive demo.
I’ve had the pleasure of sitting on a network that was, in practice, not bandwidth limited and it has led me to conclude that the terrible experience in practice is caused by retail ISPs being absolute dogshit. If you can get on a really well run ISP like Fiber7 in Switzerland, or a $BigCorp network, things are much better and demos like this are no problem.
Sure, there's post-processing, scanning your face with this level of definition may take some work currently and the full screen video we see may feel very different for someone wearing those goggles, but these should all be solved as the technology improves
You can fight it all you want, but if it's half as good IRL as this demo suggests, it's obviously here to stay.
Criticism so far seems to make the same few points:
> "Hell, now remote work is ruined. Thanks, Zuckerberg"
I actually think this may help enable remote work in the long run as companies who are unwilling to accept the current WFH/hybrid models see this as a viable compromise.
More importantly, when I call my supplier in Belgium, I'll be able to see them "face-to-face" (avatar-to-avatar?) and develop a more human relationship than just exchanging emails or phone calls.
> "Great, social relationships are going to be even worse"
If that's the use you want to make of it, sure. But it also enables you to connect with loved ones who live far away. It can be so powerful for the elderly who struggle with loneliness to feel closer to their family.
> "These two guys are terrible at conveying emotions with facial expressions"
Ad hominem notwithstanding, if anything this is an endorsement of the casting choice for a tech demo since their expressions would be easier to replicate in the virtual world
> "These two guys are terrible at conveying emotions with facial expressions"
Fridman anticipates and addresses this reaction at the 10:08 minute mark in the video as well, so for the people that watched it these type of comments will unfortunately come off as unoriginal and stale. Either the people saying that didn't get that far or they felt the need to make the comment anyway.
In several years we will be unable to say if the content is generated by an actual media person, ML agent or low-paid shadow-performer.
[1]: https://www.zdnet.com/article/meet-your-digital-persona-appl...
Now they just need to be able to "load the weapons program", and "learn kungfu".
It’s great looking tech but it couldn’t have come from a worse company.
I think it's pretty cool but a bit jarring. I think I'd rather have the cartoonish avatars until this gets a little better. It's a good start though!
And I agree that seeing facial expressions is critical for the best interactions.
To get full realism for the viewer in the headset as he moves around, they would need to be able to accomplish something like that.
We need to invent a new word for this sensation.
I'd wager that's more a product of technological limitations (and overall awkwardness) than a matter of demographics. Video games and other fully 3D environments tend to avoid photorealism at all cost, because it's compute-expensive and ugly. By comparison, simple cartoon characters, blobs or robots are inoffensive and perfectly usable abstractions. Even this "Avatar Encoder" is 'cheating' by only rendering a relatively static portion of your face. It would be almost unusable in a VRChat-style environment where dynamic lighting and shadows are concerned.
Lex Fridman is a Russian-American computer scientist, podcaster, and writer. He is an artificial intelligence researcher at the Massachusetts Institute of Technology, and hosts the Lex Fridman Podcast, a podcast and YouTube series.
The debate has already been had so i'll just link it
You were meant to chuckle at it, not take it seriously.
Congressional hearings are purely for political grandstanding. The low height seat countered with a cushion, the dumb questions answered with direct unemotional answers. 'we sell ads senator'. The entire process had nothing come out of it except a few politicians had egg on their face.
As mentioned in the video by Lex, it's the subtleties that make all the difference. I'm astonished with the accuracy of the blinking, mouth movements, subtle cheek variations, etc. It seems more accurate than the realtime feed from my webcam. The only thing I wouldn't like about it is having to wear a headset in order to experience it.
Nothing bad could ever come of that /s
The tech is cool but I'm wary of continuing to evolve the Internet to reward people for their physical appearance rather than their intellectual contributions.