Meta Project Aria - Smart Glasses Research Kit
projectaria.com
projectaria.com
In Project Aria video, they claim to have installed beacons at an airport to enable indoor location, only to dismiss it as something that "doesn't scale."
Instead, they say they "trained" an AI model using vision from glasses, allowing for vision-based localization.
So, here’s an honest question: which approach is actually easier, more cost-effective, and energy-efficient?
1) Deploying 100 or even 1,000 wireless, battery-operated beacons that last 5–7 years—something a non-tech person can set up in a day or two.
2) Training an AI model for each airport, then constantly burning compute power from camera-equipped glasses or phones that barely last a few hours.
Thoughts?
Really it's more like three questions.
1. Easier? I guess that depends how you define ease, but it largely depends on what resources you have available to you. If I'm Meta and I already have a ton of compute and AI training expertise but don't have relationships with all of the airports, stadiums, etc., their approach is probably easier. You'd have to spin up new teams of people all over the world to get beacons everywhere you want them.
2. Cost-effective? I don't know enough about the costs of your solution to give an accurate answer here, but again it just seems like they're probably already spending resources training models on a huge number of images of the world, so maybe not a lot of incremental cost here.
3. Cost efficient? I would assume your approach wins here.
It's not a big problem if you want to equip one venue or a couple, but scaling the operation means these side-effects scale too, and we had to work on solutions to handle those, rather than working on our core competency of mesh wi-fi. Unsurprisingly the project was scrapped despite being technically feasible on a small scale - we had a couple of sites.
Virtualizing a physical space gives you more flexibility. It keeps most problems in the software engineering space and limits physical requirements (eg someone might still need to walk around an airport to update the model, but I can't think of any other major ones).
That said, AI is sexy (right now), Meta is heavy in the MR space and the tech is reusable, even if it's not the most energy-efficient solution.
(disclaimer: just my personal ramblings, I don't work on project Aria)
A large part of the variation I found was due to how individual users held their phones, and the resulting signal attenuation.
[1] https://hackaday.com/2015/12/18/immersive-theatre-via-ibeaco...
In over a decade of indoor robotics I have _never_ seen a beacon-based solution that practically scales (even marker-based solutions are challenging). And it's not because the tech is even bad – it's just that any process that involves _installing things_ is a PITA and wildly more expensive and time-consuming that it should be.
This kind of sucks but it is an unfortunately reality, in my experience at least.
This is where I think the gap is. We only have to train one model for all environments, not per
How do the costs compare for training one big model vs installing billions of beacons?
Also consider the pace at which model sizes, training, and operating costs are falling
With glasses, you can map the space while identifying POIs—-with beacons you can’t.
Unfortunately, no one really cares about energy use.
Also exist some nuances, as some cities are flat but others have large hills, so need to place few beacons on sides of hill (rough surface need much more beacons).
Practically, I have experience in project to deploy LoraWan network in large city Kiev, and one concurrent bought research from cellular network planners and for first look they drawn ~300 access points to have more than 99% coverage.
2) Is a $1-??? business requiring a few dedicated nerds working on CV with inf more applications and doesn't require "invading" physical buildings you dont own.
The headset needs the inside out tracking anyway to draw specially locked virtual objects.
To create a fresh spatial anchor at home on mobile hardware is maybe 1 second of compute time. But that doesn't really matter because the anchors would be shared across every user and computed offline beforehand.
As far as scaling the device itself can be used to crowdsource these anchors so it's not even close that the visual solution wins out.
That said beacons are probably better for supporting handset platforms. Powering up modern cell phone cameras to use AR is pretty slow and tedious for the user.
The data will all be relative to initial positions and it will have drifts, but how those affect your research goals will be use case dependent(esp. since this is pitched for researches than as ready to go entertainment).
Anything you want to track in the meat realm, especially a place like an airport, the airtag or google equivalent mesh networks are going to be far more dense than your beacons and last forever with no power required.
It's their purpose in VR/AR to have cheap indoor location, for them it's one more step in that direction. Eventually they will achieve doing it with little compute.
Well, Meta poured a shit ton of money into making Quest base station free and they got there. We use to use valve setup for our robotics applications but we swapped it out with Quest cause honestly Quest was as good but much more easy to setup and operate.
The bitter lesson is that don’t bet against data or compute. Also, I don’t think you’d have to train a AI model for each location at every time in the future. Things get more efficient, etc.
I think you are asking the wrong question. The right question is: "Which approach will people use?"
Doesn't matter if it is the easiest cheapest most energy efficient thing, if people don't use it.
There are many single airports with more than 100 points of interest. Now extend that to every US state...
If you're interested in where they're currently focusing their spending and the timelines for return on investment, the recently leaked memo isn't a bad place to start https://www.uploadvr.com/meta-cto-to-staff-leaked-memo-2025-...
The Quest is a marvel and they seem to be making real gains towards a mythical hands-free AR glasses experience.
I was also very excited to see a big tech company moved towards premium hardware and premium software.
Sadly I fear a return to freemium, now AI generated, and soon to be advertisement filled slop. Meta is still Meta but hopefully their goals keep them on a better path.
He himself put Meta in that hole, found out he couldn’t manifest better hardware out of the magical money hat, and finally gave up on the idea, firing a whole bunch of engineers in the process.
[0] https://facebookresearch.github.io/projectaria_tools/docs/te...
Haha I understand why but my only real complaint about the glasses is that I'm stuck with Meta AI. Would be so nice if I could plug Gemini or Open AI into it.
Great product overall but suffers from not having an SDK and lock in on the model.
It's not a new headset or a protoype for one.
"Egocentric, multi-modal data as available on future augmented reality (AR) devices provides unique challenges and opportunities for machine perception. These future devices will need to be all-day wearable in a socially acceptable form-factor to support always available, context-aware and personalized AI applications. Our team at Meta Reality Labs Research built the Aria device, an egocentric, multi-modal data recording and streaming device with the goal to foster and accelerate research in this area. In this paper, we describe the Aria device hardware including its sensor configuration and the corresponding software tools that enable recording and processing of such data."
https://facebookresearch.github.io/projectaria_tools/docs/te...
VR and AR devices so far always use a fixed focal plane for everything. Usually around 1 meter. So, if you are looking at a distant object in VR or through AR video passthrough, your eyes need to focus at around 1 meter.
I know some VR headsets that offer customizable lens. But, I don't know about this device in particular.
I mean, that really isn't true. There have been wearable and carryable hidden cameras for ages and we also have 360 cameras that no longer need to be pointed at what they are capturing.
This isn't changing anything about what is available to purchase, and if anything, these are relatively more obvious.
The actual change would be, that if these become widely adopted, those types of cameras would be everywhere.
Historically, body-worn hidden cameras have been for perverts, spies and journalists. Normal folks don't mind people knowing they're taking a photo, and want to be able to frame the photo and suchlike. Gopros would be clearly visible, front and centre on people's helmets - and only while doing sports. Guards and cops with body cameras want people to know they've got a camera, as a deterrent.
You'd occasionally see hidden camera footage used by investigative journalists - but outside of that, the market for body-worn hidden cameras was mostly weird lonely pervs who wanted to take photos at the topless beach and upskirt photos without getting into trouble.
A glasses-camera product won't succeed among us normal folk if wearing it makes you look like a weird lonely perv.
You say this, and yet as a blind user of the Meta glasses (they're actually great for accessibility!) I am not ... seeing it. They are far more ubiquitous and warn by far more people than you would expect, especially when comparing to Google Glass.
If this were pushed by Apple, people would be responding much differently, since there is an inherent level of trust in regards to Apple's privacy protections, vs Meta.
So, not so much the technology, but rather the trust (or lack thereof) behind the implementor.