FTR: https://en.wikipedia.org/wiki/Dragon_NaturallySpeaking
Dragon Systems released NaturallySpeaking 1.0 as their first continuous dictation product in 1997. <...> As of 2012 LG Smart TVs include voice recognition feature powered by the same speech engine as Dragon NaturallySpeaking.
Today voice transcription is a solved problem and while their engine might be the same in name - I’d be surprised if the approach isn’t totally different than what they were doing in 97, either that or the LG tv voice transcription probably doesn’t work as well as everyone else’s.
The deep learning revolution and the applications we’ve seen since 2015 are a major step forward and something truly different. People pretending otherwise are just acting cynical in some attempt to project intelligence or seem wise, it doesn’t work.
But you can't claim that something "wasn't possible 5 years" ago, if 7 years ago said feature was included in inexpensive consumer product (LG TV).
I'm not acting cynical, but it's tiresome for me to see people who claim that 20-30 years ago we all were living in a caves and catching bugs with wooden sticks, and now boom, ML!
Regarding "something truly different", well, my personal computing / mobile experience not changed that much from 2015. Honestly speaking, progress from 1995 to 2000 felt much more impressive and 'truly different'. I mean, think of it, during this timeframe we went from DOOM via V.34 modems to amazon.com and ordering pizza online.
I feel like an AGI could accidentally wipe out half of humanity and there would still be people commenting on HN about how the exact same technology already existed in a roomba seven years ago.
- Voice transcription
- Tesla Autopilot
- Facial recognition (photo sorting on iphones, better photos)
- Better graphical performance on Nvidia cards (https://developer.nvidia.com/dlss), also better compression for streaming.
- Much better translation
- Colorizing and repairing old photos
- Visual recognition allowing better search of images
I’m sure there are some I left out. I think we’ll see a lot more interesting applications (particularly around tooling) in the next few years.
https://medium.com/@karpathy/software-2-0-a64152b37c35
Outside of the consumer space, there are also things that hint at more generalizable intelligence.
Check out GPT-3’s performance on arithmetic tasks in the original paper (https://arxiv.org/abs/2005.14165)
Pages: 21-23, 63
Which shows some generality, the best way to accurately predict an arithmetic answer is to deduce how the mathematical rules work. That paper shows some evidence of that and that’s just from a relatively dumb predict what comes next model.
It’s hard to predict timelines for this kind of thing, and people are notoriously bad at it. Nobody would have predicted the results we’re seeing today in 2010. What would you expect to see in the years leading up to AGI? Does what we’re seeing look like failure?
https://deepmind.com/blog/article/muzero-mastering-go-chess-...
There have been massive improvements in automated driving, but if you want to talk about solved problems, parking assisst is as far as you can get.
Translation is much better, and is often understandable, but it is far from a solved problem.
Colorizing/repairing old photos also often introduces strange artifacts in places where they are unnecessary. Again, workable technology, not a solved problem.
Voice transcription is also decent, but far from a solved problem. You need only look at YouTube auto-generated captions to see both how far it has come and how many trivial errors it still has.
And regarding "generalizable intelligence" and arithmetic in GPT-3, the paper can't even definitively confirm that the examples that they showed are not part of the corpus (they note that they made some attempts to find them that didn't turn out anything, but they can't go so far as to say they are certain that the particular calculations were not found within the corpus). They also make no attempts to check the model itself to find if any sub-structure may have simply encoded an addition table for 2-digit numbers.
Also, AGI will certainly require at least some attempts to get models to learn about the rules of the real world, the myriad bits of knowledge that we are born with that are not normally captured in any kinds of text you might train your AI on (the idea of objects and object permanence, the intelligent agent model of the world, the mechanical interaction model of the world etc.).
'Solved' is doing a lot of work here and is an unnecessary threshold I'm willing to concede. I think we would be more likely to agree that things went from unusably bad or impossible to usably good, but imperfect in a lot of consumer categories in the last five years due to ML and deep learning approaches.
The more clearly 'solved' cases of previously open problems (Go, protein folding, facial recognition, etc.) are mostly not in the consumer space.
As far as the gpt-3 bit I encourage others to read that excerpt, they explicitly state they excluded the problems they asked from the test data so it's not memorization. The types of failures it makes are failures like failing to carry the one, it certainly seems like it's deducing the rules. It'll be interesting to see what happens in gpt-4 as they continue to scale it up.
Even the good translation services will randomly produce garbage, struggle with content that's not in nicely fully formed sentences, and are severely limited in language support.
Similarly, both Amazon [1] and online pizza ordering [2] existed before 1995. They were just not commonly used.
[1] https://en.wikipedia.org/wiki/Amazon_(company)
[2] https://www.zakon.org/robert/internet/timeline/index.html
Siri from Apple was launched in 2011, as some other commenter noted below. Also, "On June 14, 2011, Google announced at its Inside Google Search event that it would start to roll out Voice Search on Google.com during the coming days".
If it does not count as 'typical consumer to experience it', well, I do not know what counts then.
9 years ago, I mind you, not 5. And I think that 5 years ago voice recognition was more-or-less good already. In 4 years both Apple and Google acquired large enough datasets to learn from, afer initial launch of their products in 2011.
What we are still struggling with is proccesing of fuzzy queries, something among the lines of 'Siri tell me which restaurant in my area serves the most delicious sushi according to yelp reviews and also allows takeout', but this is not a voice recognition problem (though typical consumer can think it is).
Siri stumbles at way less complex queries than that. Every year or so I retry using it, and give up due to the error rate. An accuracy of 99% and 10x slower is apparently preferable for me.
I'd take a literal clapper that hooked into smartbulbs over it at this point.
In five years we will without a doubt have systems that play more games, solve more puzzles, fold more proteins and so on, but we'd have more progress on intelligence if we could build a machine that could reliably plumb a toilet or catch a mouse as well as a cat.
Voice recognition is neat, and it is extremely useful in certain niches, but it is generally far inferior to mouse+keyboard or touch or dedicated buttons in all cases where those are practical. It is almost unusable in public spaces. Text-to-speech is mostly in the same boat - extremely useful in some niches, not a game-changer in any way I expect.
Text generation is only "realistic" in shape, but for now remains a novelty. Picture generation is the same. They are at best at the level where it already costs pennies to obtain similar texts through human labor. Having conversations in a meaningful way is even farther.
Playing Go better than humans is a fun gimmick - an awesome achievement in some sense, but with 0 practical applications outside of Go.
Overall, there has been good progress in AI/ML, but I don't expect we'll see any major changes too soon from this. Especially since the idea of actually seeing commercially available self-driving cars has been pushed back a decade or three compared to the initial optimism.
I'd also add that most of the progress in ML/AI in the last few years has essentially been of a technological/engineering kind. There haven't really been that many deep scientific insights, we don't necessarily understand any of these problem spaces better. We've just been able to match certain neural net architectures with certain problem spaces, and massively advanced in realistically-available computing power to actually train these nets on enough data.
But we haven't gotten any new insights into stuff like language from GPT-3, for example. We haven't even gotten any new insights into Go from AlphaGo!
I don't care about driving. It still takes the same amount of time.
What I care more about is cooking. I'd rather have that one solved, as it would actually save me time. (No, I'm not looking for restaurants, meal delivery services or microwave meals).
If it would be automated, the you would be able to use this time for something else rather than driving.
AFAIK the use of learned embedding in the biomedical space has improved our understanding of the science/our ability to find high value experiments to run.
I think the coherency of GPT-3 and the RL learned tool use work both raise extremely interesting questions about the nature of language and high-level intelligence. Questions that we couldn’t meaningfully ask before that evidence.
I also think you’re being shortsighted about the possibilities that are opened up as we enable computers to understand and interact with the real world in more way. Increased visual and audio understanding will lead to more passive computing and that has all sorts of fascinating implications.
I don’t know what insights you’re looking for. We’re seeing AI reveal fascinating phenomena in language, vision and general purpose learning. Scientific advancement is often about asking the right questions and unexplainable phenomena is the driver of those questions. Much like some phenomena (I want to say something to do spectral) were unanswered questions that developed into the field of quantum mechanics, what we’re seeing in AI is currently unexplainable phenomena.
Are you in the field? You’ve focused on the things that made global headlines, but I feel like there’s so much more.
We've learned how to play better Go, but not anything like a new (mathematical) theory of Go, as far as I have read. And I didn't claim it's just about computing speed improvements, I also noted that we have started the right formulas to fit specific NN architectures with specific kinds of problems.
> AFAIK the use of learned embedding in the biomedical space has improved our understanding of the science/our ability to find high value experiments to run.
Awesome, this is one thing I hadn't heard about.
> I think the coherency of GPT-3 and the RL learned tool use work both raise extremely interesting questions about the nature of language and high-level intelligence. Questions that we couldn’t meaningfully ask before that evidence.
They do raise interesting questions in a philosophical sense, but those questions do not seem to have piqued the interest of the AI community. The GPT-3 paper covers some interesting facts about GPT-3's ability to do arithmetic, but only briefly, and that's about it. They explicitly attribute GPT-3's impressive gains over GPT-2 entirely to the hugely increased parameter space, and mention that they believe that increasing the parameter space again by an order of magnitude will increase the realness even more - that's the extent of their analysis about its implications on language that I've seen (perhaps I missed something?).
I haven't even seen a real discussion about how different GPT-3's output is compared to it's training inputs. Given the gigantic corpus that they have used to train it, I would expect to see more scrutiny in this area. For the arithmetic operations, they mention some attempts they made to check that it wasn't reproducing a calculation it had explicitly seen, but they are not confident that they were enough.
More importantly, it is obvious that, even if GPT-{X} will be a complete model of human language, it will (1) have nothing to do with how humans acquired language, either individually or evolutionarily; and (2) that it won't be able to produce meaning that is not captured in its corpus, as it will have no notion of the human world and its reality; it might be able to produce commentary on current events that sounds plausible, but any relation between its commentary and reality will be either entirely captured in the training corpus, in the priming text, or entirely accidental.
> I also think you’re being shortsighted about the possibilities that are opened up as we enable computers to understand and interact with the real world in more way. Increased visual and audio understanding will lead to more passive computing and that has all sorts of fascinating implications.
Of course that's a strong possibility - the future is usually surprising. However, the only applications actually visible on the horizon are bone-chilling: mass surveillance that even Orwell didn't dream of, increasingly being actively used by more and more authoritarian regimes, from China to the US to Europe. It may well soon turn out that the AI fearmongers were right to fear AI, though of course not in the puerile fantasy of the paper-clip optimizer.
> We’re seeing AI reveal fascinating phenomena in language, vision and general purpose learning.
I have seen little to no coverage of such fascinating phenomena, and entirely too much coverage about precision rates, eerily human-like language generation, and a belief that bigger models and more data are the only way forward. Everything I have seen has been AI research not only not asking questions about such phenomena, but instead shutting down any questions, claiming that the million-parameter models that they produce on terrabytes of training ARE the answers [0].
> Are you in the field? You’ve focused on the things that made global headlines, but I feel like there’s so much more.
I am not in the field of AI/ML, no (the closest I got was doing my bachelor's thesis on an RL approach). The things I focused on were the advances highlighted by the post I replied to. I have no doubts that there are many advances in the techniques of AI that would completely fly over my head, and I am certain that there are successful applications of AI on hard problems that I never even heard about. AI is certainly a useful technique in many fields, though often over-hyped as well. Still, judging by what is commonly highlighted as the major achievements of the field, my prediction is still that it's positive impact on the world will continue to be limitted for the next 1-2 decades; its negative impact through enabling mass surveillance may well be far greater.
I’m not saying there hasn’t been innovation in the past 5 years — absolutely there has.
In that industry, 2013-2016 were peak “ML” and “Cloud” where the CIOs at my company and competitors were fully bought into the ML and cloud hype, how it was going to solve all kinds of problems, without understanding what those problems were, and without realizing the complexity of getting meaningful, applicable data for those problems.
In the past few years, it feels like more people realized that ML is less mathematical magic and more of different “kinds” of curve fitting.