Amazon's "Just Walk Out" checkout turns out to be 1000 workers watching you shop
boingboing.net
boingboing.net
Outsourcing to India/Mechanical Turk is always going to be cheaper than creating a sophisticated AI model. I could see them dropping it if the "learning" had started to stall.
This whole thing seemed like a branding exercise anyway. Like how Apple stores are in every major city despite relatively few people buying their products in-store.
(With apologies to Matt Levine)
Just don’t act surprised when we “find out” the dash carts are also powered by humans
Ultimately it might be as simple as the fact that even cheap ass labor plus the infrastructure plus the AI itself still costs something and ultimately that something is more than just paying some friggin cashiers at the bloody store to do the job.
My first question is what does a GPT have to do with solving this problem? Honest question, because I don't see how a language model helps solve this, so if your do have a clear explanation it will be appreciated. I get they parse image data but have they demonstrated better accuracy for this use case (grocery items/food) than what was available when Amazon started their campaign?
There's also the fact that, yes more data gives better results... To a point. At this stage, I'm inclined to think we've hit that point when it comes to deep learning models' ability for object recognition and we're in a plateau until a rational new approach (i.e. not deep learning) allows for another leap.
My last is more socialogically rooted. Namely, it's pretty clear that since sama said he'd need $7 Trillion, AI peddlers have jumped the shark and we're seeing a slow motion deflation of the major claims, even concerning LLMs (don't get me wrong, very impressive tech, but no where near the heights and capabilities of the hucksters). For that reason I'm very inclined to lean towards skepticism when people claim they've found a solution by incremental improvements to a problem that's proven to be very, very difficult for a computer to handle.
I think we're nearing a general plateau in deep learning and need the next foundational approach or what we could find ourselves in another deep AI winter.
Any reason you think my skepticism is unfounded and why 5 years?
Have you seen how effective transformer-based networks are at image recognition? They are literally better at reading bad cursive handwriting than I am, at this point. So yes, they can be pressed into service recognizing products being taken off of shelves and placed into carts or bags.
"Just walk out" is a difficult problem but very much worth solving. We now have tools to attack it that nobody was even dreaming of when Amazon designed their current approach. Tools with near science-fiction levels of capability.
Is there a specific benchmark comparison between a GPT and any OCR app/algo? Any cases where you've seen a free OCR app fail spectacularly where a GPT succeeded without question?
Language models are just that - language models. They are good at predicting what word should come next. Because it's language we impart capabilities way above and beyond their reality. Sure, they can incorporate CV data (among other inputs) that feed them the resulting text. Maybe when we develop something with a functional world model, we might be able to tackle the real physical world. Until then, problems like this, self driving, folding clothes, etc will keep failing at the edge cases, rendering them more trouble than they are worth.
And "just walk out" is a fun, interesting problem only in that it might give rise to interesting solutions and approaches for more pressing or relevant problems
Go ahead, tell me. Life is short of opportunities for good sensible chuckles.
It tells me this is a solution without a problem.
I'm not saying the feat of LLMs is not impressive, they certainly are. Just don't tell me they have developed a "world model" and display understanding because they have not. They will always suffer the same issues that have plagued self-driving cars and autonomous robotics: they can only process what's in their training data, therefore they need to be trained on all scenarios that will ever exist to function outside of well-curated, well-defined closed systems.
I would love a good chuckle too, unfortunately the total lack of critical thinking and understanding when it comes to these stochastic correlative black boxes leaves me greatly disappointed
[1] https://www.reddit.com/r/Wellthatsucks/comments/j67atm/1_sec...
More discussion: https://news.ycombinator.com/item?id=39908579
Let's zoom in on Lahore, pulsating with life and a fervor for innovation that makes Silicon Valley look like a retirement home for tired tech. Here, in a place that hums with the energy of endless possibility (and perhaps too much caffeine), sits the command center of the world's most "advanced" computer vision system, SeeAll. Only, SeeAll's vision is purely human, powered by a legion of sharp-eyed annotators who scrutinize live feeds from grocery store checkouts, identifying every item with the accuracy and flair only a human can muster.
The masterminds behind this grand ruse? VisioTech, a company shrouded in the mystique of technological advancement, promising a checkout experience free from human error, barcodes, and the tedium of waiting. Their secret sauce wasn't algorithmic; it was organic, brainpower fueled by chai and an undying spirit of camaraderie.
This narrative, though dripping with the trappings of high-tech, is really an ode to the human element in the digital age. Customers, awash in the glow of seamless transactions, never questioned the "how" of their flawless shopping experience, unaware that the real magic lay in the hands of the PaaS team. This crew, capable of identifying the most obscure products with a glance, operated in a symphony of clicks and keystrokes, a ballet of productivity that no machine could replicate.
The facade crumbled when a technical hiccup at an Ohio grocery store exposed the gears behind the magic. What followed was a tumultuous unraveling that laid bare VisioTech's elaborate scheme. Yet, the fallout was anything but predictable. The world, instead of recoiling in outrage, leaned in, captivated by the audacity and sheer inventiveness of it all.
The only way this can be interpreted as "people manually look at the camera and tabulate the price" is disingenuously. And it's one example that I suspect will be scrubbed from the web, but from the moment this was launched, unless you were living under a rock, the message was "AI powers the store." So please stop with the hairsplitting and revisionist history.
(1) https://aws.amazon.com/blogs/industries/enhancing-the-retail...
E.g. the way I imagine it to be is that this CV system will attempt to identify the item and confidence that it is that item. If the confidence is too low, it will trigger a signal for human feedback, then human will either confirm or reject.