What Happened in the 2010s
avc.com
avc.com
What are some examples of companies that hinged on machine learning?
DATA is crucial, but that's very different than "sophistiated machine learning models". For example, do Uber and Lyft need machine learning?
As far as I can tell, they need data about drivers, users, ride completion, maps, and a lot of people doing simple statistics on that data and improving the product. Machine learning might give them a few percentage points in some specific subsystem, but it isn't going to make or break the product as a whole. It's not "table stakes".
Pick another industry and do the same thought experiment. What are some other breakout companies? Did snapchat succeed (in its day) because of machine learning? What about AirBNB? I think they are optimizing the product with "data science", not machine learning. Certainly not deep learning.
Machine learning finally came of age in the 2010s and is now table stakes for every tech company, large and small. Accumulating a data asset around your product and service and using sophisticated machine learning models to personalize and improve your product is not a nice to have.
My favorite is that they do pothole detection, prediction and routing around based on what they classified based on driver accelerometer data. They've also reduced traffic incidents with insights from the same dataset that drove product changes.
To me, the accelerometer DATA is the new and important thing -- it's the thing that taxi companies don't have.
Whether the machine learning model is "sophisticated" doesn't seem critical on the face of it.
I guess it depends whether there are a lot of non-pothole events that look like pothole events to "naive" signal processing. Do you need deep learning for this problem? I don't know but I would be skeptical without a citation.
But I am interested in more info and examples. If someone can give a bunch of examples I would adjust my view.
Data is obviously crucial, but it's useless if it's just sitting on hard drive. Machine learning allows them to squeeze out as much juice from it as possible, and a few % point lift makes a huge difference at scale, especially in high volume businesses with low margins per transaction.
More importantly it's worth understanding that ML allows us to do things we couldn't do before, because it's cheaper and scales better than previous optimization methods, when you have the data.
So yes, the data infrastructure is the critical part, however the ML pipelines and Data Science work is what gives the massive multiplier at a cost/scale you can't get with traditional optimization methods.
Alexa / Siri / Google Assistant would not be possible without machine learning. Same goes for Snapchat's filters and any other form of AR. Google's Pixel phones wouldn't have such a great camera without computational photography / machine learning.
That said, uber and lyft certainly benefit from much more interesting ML and stats than just "simple statistics." For example, setting prices and incentives to properly balance the driver side against the rider side of the marketplace.
And yes I probably should have said "data science" or "statistics" rather than "simple statistics", definitely for Uber and Lyft (less so for smaller companies).
I worked on a data science team 10 years ago when it was called "statistics". I even remember a 2009 NYTimes article that talked about data science without using the word "data science", because it hadn't come into use yet.
If people want to redefine "machine learning" to mean "data science", OK fine. I guess that's what's happening now. The term "data science" already annoyed some people, but I guess now it's "machine learning" and "deep learning".
Anyone doing large scale ETA prediction and routing these days is using ML models at various phases in the pipeline, from GPS noise filtering to road network data generation. You might get past the early stages of a rideshare or delivery service with more traditional naive models, but there are still significant efficiency gains to be made in even the most cutting edge systems today. For maps and navigation, this is especially true when you go global and need models that adapt to subtle cultural and behavioral changes across markets.
In the next decade it will start taking over computer graphics, medicine, manufacturing, surveillance, hardware / software and every other aspect of our lives.
EDIT: for people who are downvoting me, here are some examples of how machine learning is and will transform our lives:
1. 3D graphics: https://www.youtube.com/watch?v=FlgLxSLsYWQ Andrew Price doesn't have a technical background but does a decent job summarizing some of the ways deep learning will transform computer graphics. Soon all of the mocap, character design and animation will be driven by deep learning systems.
2. Computer Architecture and Traditional Software: https://www.youtube.com/watch?v=TTpKWOuzOxc There's a ton of recent research showing that you can use machine learning to beat human crafted heuristics in hardware, scheduling, compiler design and query planning.
I would say it's table stakes for large companies in certain areas, but not small (see my other reply).
Machine learning seems to be good at squeezing percentage points out of monopoly with a business that already works. That's very different than "table stakes for small companies". Small companies need optimize their product with fundamental changes.
For example, in the early days of AirBNB there was a more inconsistent user experience, which you heard about in the media. I think they addressed those problems with policies, incentives, and some elbow grease (kicking bad users off the platform, etc.). Not machine learning. Machine learning doesn't help you expand into Europe or Asia, etc.
Yep, and the results are hilarious or tragic depending on how you look at it:
We keep getting ads for the thing we bought yesterday, ads for dating sites after we got married and had kids (continously, for ten years despite my utter lack of interest), search results keep getting worse[0], obvious spammers keep on spamming in social media (seriously, it seems a simple regex filter could have done a better job to reduce crypto scamming in replies on Twitter than whatever was there last time I checked.)
[0]: some people will always claim it is because black hat SEO is so much worse, but that doesn't explain why Google sometimes can neither understand doublequotes nor the verbatim option anymore. That is not the result of black hat SEO but of sloppy maintenance, and I guess so is a number of other problems.
That's because in a lot of cases they're optimizing for the wrong metrics, as in maximizing their revenue instead of your utility.
There's way more content on the web than there was in the early 2000s, most of it in form of "content marketing" and explicitly attempting to game the system.
If you look at recent results from TREC, it's pretty clear that machine learning provides a large boost over the traditional retrieval systems, on any metric that you want to optimize.
What? How do they maximize their revenue by spending money on ads the users actually laugh about for how bad their targeting is?
Based on what I know about machine learning - it almost always gives great short-term results, and it almost always fails to deliver the expected long-term results. What's worse is this comes with the weirdest most indecipherable bugs that pop up more and more over time. Unless you have a large enough database to show statistical errors that are negligible (something like 99.9999% precision or recall, depending on your metrics), you should assume it will break in ways you cannot possibly predict. And even then, you might be using the wrong training data without even realizing it.
I'm not saying ML is bad, although I am saying it is ridiculously overhyped. I'm saying ML is still nascent enough nobody really knows how a lot of edge cases will shake out, simply because there are too many edge cases to test before putting it into production.
It's not hard to find examples of ML algorithms gone wrong even for sites like Amazon.
https://gizmodo.com/amazon-prime-day-glitch-let-people-buy-1...
You can have a system that generates them a ton of cash while making mistakes in some cases, outliers are inevitable and feedback loops in recommender systems can lead to such issues. Amazon wouldn't deploy these systems if they didn't move the needle.
> Amazon wouldn't deploy these systems if they didn't move the needle.
This is an appeal to authority that Amazon executives are immune to making mistakes. They could have bad metrics. They could have bad incentives encouraging managers to make poor decisions (basically this describes everything wrong with Google today). They could be incompetent. They could be focused on short-term gains at the expense of long-term gains because it maximizes their personal net wealth and they can just jump ship in a few years.
Again, nobody is saying ML algorithms are worthless. I just don't believe (and have lots of reasons based on personal experience I won't go into) that they are 10% as useful as the industry wants you to believe.
Anyway, agree to disagree.
There's a story, and I think it was (re)posted here recently about a series of MBAs all optimizing the cost of the burger bums by removing seeds until there are three seeds neatly laid out at the top.
It might be measurable all the way but at some point it becomes ridiculous. For me that time was some months ago. For the rest of Internet they might manage to reduce the quality once or twice more before it becomes obvious.
This is my way of fighting back. By posting here and on my blog and getting upvoted a lot for pointing out what many can already feel. By letting people who read HN know that yes, that feeling they have that a lot of their ads are wasted because of bad targeting might very well be true.
I'm against most forms of recommendations as well, that doesn't mean that they're not valuable to the businesses deploying them. In most cases the end users of these systems are not the real customers of these platforms, and they're there to serve the advertisers who keep these businesses running.
Not if it's something I only buy once a year, for example. That's where "learning" part should come in. You don't need any learning to just parrot me back my inputs.
> Amazon wouldn't deploy these systems if they didn't move the needle.
I don't know about Amazon, but I've recently read on HN some articles strongly suggesting almost nobody is properly measuring the impact of ads, let alone the impact of "targeting". In most cases, people more or less just stuff money into ads budgets, because that's what you do, and they get customers - because people still need to buy things, regardless of any targeting - but the casual link between the former and the latter is not really very well established.
I am not sure how these are contradictory. My interest is buying things I want. Their interest is selling me things I'd buy. How showing dating site ads to a married person promotes any of those? Where the revenue would come from? Or do you mean it's ad agency revenue, not advertiser revenue? In that case we clearly have a case of agent problem.
This does not demonstrate that the machine learning algorithms are doing a bad job. Quite the opposite really. See:
As for ad placement, social media, and search ranking: is that actually effective? The one constant with ads and social media and web search, IME, is that it just never gets any better. From the Lycos/AltaVista days to Google there was a huge jump in search result quality, and then it plateaued.
The ads I see online are still laughably bad. They're more professionally produced than 10 or 20 years ago, but they're still never for any category of product that I'd ever buy. Targeted and tracked advertising seems like a big scam.
Those were just some of the most visible examples. OCR, image recognition, NLP and speech recognition alone are enough to revolutionize most data entry jobs and we'll see a ton of startups using them to do just that for every industry.
There are countless other examples of machine learning applications, including drug discovery, radiology, production line quality assurance, etc.
> The ads I see online are still laughably bad.
Have ads ever been good? You see the ads that you're seeing because someone is paying a lot of money to put them there.
Bloomberg spent 120 million on ads last month, ad platforms had to find eyeballs for them.
If I'm looking at, for example, "how to programmatically control my model railway with a Raspberry Pi", there's a large number of easy to target, relevant products you can advertise against that. You don't need a gigabyte of creepy backstory on me to figure out "hey, show me ads for the products mentioned in the article, and there's a good chance I'll buy them because it's convenient."
Yeah, Mike Bloomberg might be willing to pay $1 to show his ugly mug next to that page, but the ad from a modeler's supply shop, who would only have paid 95 cents, would have felt much more appropriate to the consumer. I know some ad networks try to measure quality via clickthrough rate, but I'm not sure it was weighted in a way to nerf this issue.
Another problem is that a pure-content model doesn't extend to all types of site. There's no suitable product to advertise against "six die in New Year's party gone horribly wrong." and nobody wanted to run low-yield CPM branding ads, so they had to backfill wth retargeting and profile-based ads instead.
There's still a lot of anti-data sentiment in the paid ad industry, where media buyers will guide themselves by what they think their customer looks like, and not what data tells them it is.
And until this is a fixed behavior, you'll keep on seeing untargeted ads.
I have bought many things from ads related to the media I am reading. For example, I read muscle car magazines, and buy parts/tools from the ads in those magazines. Back in the era of print mags, my company would do well placing ads for compilers next to relevant programming articles.
The counter-example to your point is that deep learning maniacs were predicting self driving to be a done deal by 2020 and we now realize we are so far away from that goal. Machine learning will only work well for one-pony tricks problems that can be clearly isolated. Nothing like "every other aspect of our lives".
And the results are, as far as I can see, horrible. I mean yes, if I make a mistake to search for shoes on Google, half of the internet will be showing me shoe ads for the next 6 months (how many feet do you think I have? Do you think I buy shoes in dozens?) - but presenting it as the triumph of ML is IMHO an overreach.
> personalization on all top social/media apps
From what I see, the same apps struggle to not get sued because of the said personalization regularly pushes sex content on kids, triggers on snowflakes and election interference on potential voters. Or at least that's what I hear every time next round of censorship is introduced into the social media. Somehow I am still not seeing a cause for celebration here.
> search ranking for google
Same as above, plus one shouldn't use Google search anyway. Use DuckDuckGo.
> translation, speech recognition
OK, here it got pretty good results, though sometimes it is as good as a very drunk chimp who got into a dictionary store, but in many other times it's decent. You get this one.
> finance, fraud detection
As a consumer, haven't noticed it. 100% of fraud on my cards have been detected by me reviewing my credit card statements. 100% of fraud alerts by banks have been false positives. I do not begrudge that specifically, I'd better have false positives than more fraud, but not seeing much ML-driven progress there tbh.
> In the next decade it will start taking over computer graphics, medicine, manufacturing, surveillance, hardware / software and every other aspect of our lives.
Not sure what "computer graphics" means, medicine probably not, surveillance maybe, but that's exactly the opposite of what I'd want, software definitely not even close, for the rest I'm not even sure what you're talking about.
> here's a ton of recent research showing that you can use machine learning to beat human crafted heuristics in hardware, scheduling, compiler design and query planning.
I'll believe it when I see ML-driven code generator doing something useful without tight supervision by humans. I can believe ML doing specific heuristics better (heck, doing exactly that has been part of my job recently) but there's a huge difference between "figure out exactly how much sugar makes the specific cookie recipe taste the best" and "invent whole cookie recipe from scratch and bake the cookie". I am sure ML would be useful for the former, for the latter... I'll believe it when I see it.
Having worked in this industry for a while I just want to point out that the vast majority of fraud detection is done for the merchant, not the customer.
Pretty sure ML have shown promising results in cancer classification in images (such as xray, etc). Not sure how this will pan out but the limited scope seems ideal .
I do think in general Machine learning is only just started with respect to All industry, not just tech. There is so much shown what could be done, the next 10 years will be very interesting.
Machine Learning has made amazing strides in the past decade and will change all aspects of our lives. There will be a ton of startups in the next decade using speech recognition, image recognition, natural language processing and etc to automate all of our repetitive tasks. If hardware keeps up we'll probably see a boom in robotics in the next 5-15 years that will allow us to automate most manufacturing and food production.
AI is a branch of computer science, but outside of A* search it's practically all machine learning.
I believe the demand for companies to actually use the fancier parts of machine learning are going to be much fewer than people think. Number 2, on the other hand, is much more generalizable, and is having a much larger effect on the 'average' business than NLP and image recognition.
I used to work at an image recognition API startup and went through an AI incubator with my own company so I got to see the huge range of applications of this technology. At the incubator we had a company automating customer support (acquired by google), another one automating appointment booking and scheduling, sales call analytics using speech recognition, automatic reading comprehension quizzes for K-12, deep learning for radiology, patient tracking/monitoring at retirement homes and hospitals, imitation learning based robotics, deep learning based hedge fund, amazon go like vending machines, fashion analytics, social media analytics/engagement for brands.
I use my Google Home Mini every day and it's just like the computer in Star Trek. If you could show it to someone from the early 90s they would certainly say it resembles an AI system.
The computer in Star Trek is just a smart speaker. Lt. Cmdr. Data is AI.
I think the usual terms we use nowadays is narrow AI and general AI to differentiate both types.
Real-world win conditions tend to be more complex. For non-general AI tasks, this is much less of a problem. "Identify hot dog vs. not-hot-dog accurately 99% of the time" is measurable and verifiable. But a simple goal for a complex general-AI task is likely to build a Paperclip Optimizer. I suspect we don't even have the right language to build a good win condition for a lot of real world problems-- we may end up ignoring "insignificant" factors only to discover they're huge over tine or at scale.
Games also tend to provide near-perfect or perfect information. At a minimum, there's constrained state transitions. I can't pull out my Oyster card in the middle of a poker game, add it to my hand, and declare it's actually a four of clubs. Real problems can have surprise off-the-wall, or outright externally triggered transitions, that completely throw the model off.
"AI, build me a market-beating stock portfolio to cash out 2040-01-01" is a lot harder than "AI, checkmate this king."
I was talking about the effect on our culture, and the shift that has occurred in the last decade has changed how we consume media in a way that did not happen with what we had in the 90s.
I remember the 90s video streaming, with Real Player and the like.... it was not a viable platform for watching tv quality video... the quality was awful, and it would take minutes before you could start playing a thirty second video.
You can store a whopping five 4K equivalent movies on your 512gb SD card. You can shoehorn 10 to 20 movies on there at approximate Blu-ray quality. If you're willing to sink to Netflix equivalent pseudo HD quality you can get 100.
The piracy genie is still being very strongly held back by the massive size of the files involved. The average consumer isn't going to buy and manage ten hard drives to store everything. They're not going to spend the thousand dollars on the drives, and they're also not spending a couple thousand dollars in blown bandwidth caps to download it all. There is no scenario where the average consumer has any interest in that scenario, they'll always choose Netflix + Disney + Prime for sub $30 / month. It's a no-brainer for them.
Also, in the not very distant future those lower quality movies - at eg 5gb file sizes - will noticeably look like trash on modern TVs.
It's only true about music.
Adaptive streaming originated as a Microsoft pilot project for the Beijing 2008 olympics. Microsoft's prototype had so many hacks and so little documentation that they publically admitted to giving up on standardizing it and threw their support behind mpeg DASH in ~2011.
Also, YouTube reached profitability in 2015 when it was ported to Google's lowest-TCO infrastructure. YouTube monetization practically invented the 10 minute video and made millionaires out of many online personalities.
2010s, Smartphone ( iPhone 3GS at the time ) went from niche to 4B users ( iOS, Androids and KaiOS ), that is nearly every person on earth above age 14 in developed countries. I dont think there has ever been a product or technology innovation as important that spread faster than Smartphone. And it changes everyone's life. The post mentioned of Google, Facebook and Amazon empire, all partly grows to this point because of Smartphones. Technology companies together now worth close to 10 Trillions. The whole manufacturing supply chain exists and became huge in Shenzhen because of Smartphone. It was the reason why TSMC managed to catch up to Intel in both capacity and leading edge node. It was the reason why we went from 3G to 5G in mere 10 years because of all the investment kept pouring it. It was the reason why everyone went on to the Internet and had Internet economy. It brought a handheld PC and Internet to a much wider audience.
I would even argue it was the Smartphone innovation that saved us from the post 2008 Financial Crisis doom as it created so much wealth, innovation and opportunities.
And the 2nd most important thing not mentioned in the article. ( May be it is not important to others... )
We lost the man who bought us the Smartphone era; Steve Jobs.
The most significant negative aspect that subscriptions avoid is ephemeral spambot accounts. If a user account costs money, there's way more disincentive for behavior that will get that account banned.
The transit situation in particular will not improve because people prefer cars. People have little experience with public transit, and the experience they do have was bad, so they don't see themselves using it on a regular basis. They want to live in suburbs and drive cars, and they will try very hard to prevent anyone from increasing the density of housing or transit anywhere near their houses.
Those people vote, so a state government that tries to push for more public transit and housing density would not do well in the next election.
“Transit” and “Density” isn’t the problem in the Bay Area — it’s that it’s excessively punitive to attempt to build anything there. If you were to be able to afford available land, cities such as Mountain View want to milk developers to pay for things that the city should be paying for themselves out of normal tax revenue. I remember when Steve Jobs wanted to build Apple Park and some fool on city council wanted Apple to provide WiFi for the city for free: Jobs said that it isn’t Apple’s job to provide city services — Apple pays their property taxes, if Cupertino wanted free WiFi, then Cupertino should do it. There are also cases where if a developer wanted to build some houses; they’d also have to give the city a free park and dedicate a certain number of houses for “low income.” That isn’t a developer’s problem. If the city wants low income housing, the city should use some of its own land and budget to build it. You don’t have that sort of extortion happening in places like Houston — and coincidentally, you don’t have affordability problems in Houston either. A developer also is shy about investing hundreds of millions into a project only to get it derailed by NIMBYs or other activists or worse, be a single election away from having severe rent control or confiscatory tax policies enacted that would destroy any hope of a profit from the substantial risk of property development.
I don’t disagree with allowing high density where it is appropriate, but density isn’t the issue, it’s the extreme high costs, both political and monetary of doing business in the region. There is plenty of land, but given than almost every empty cow pasture ends up getting claimed as “open space,” makes it really hard to build anything. Case in point: the “preservation” of North Coyote Valley. I appreciate open space, however, they literally preserved 900 acres of flat farmland to be used for hiking despite the region being chock full of thousands of acres of open space already. You can’t claim a housing shortage while simultaneously taking away 900 prime acres that would be perfect for housing developments. Clearly there isn’t a housing problem when so much easily buildable land can be sequestered away so vegan hipsters can have yet another thousand acres to roam with their rescue dogs. The Bay Area has no shortage of open space, yet they keep adding more to it while people literally sleep in RVs in El Camino and a young family has no hope of ever owning a house in the area.
Edit: ok read the rest.
- according to his admission, the subscription model did not scale
- missing is the effect of planet-level monetary policy that massively benefited shares and their largest owners
> 1/ The emergence of the big four web/mobile monopolies; Apple, Google, Amazon, and Facebook
I'm not sure you could class Apple alongside some of the others for monopoly power either. They're still less popular than Microsoft for desktop operating systems, Google for mobile ones and virtually every other major service in their other areas of business.
Apple is a very profitable business, but they're arguably more of a luxury goods maker than a monopoly right now, especially outside of tech circles.
> 2/ The massive experiment in using capital as a moat to build startups into sustainable businesses has now played out and we can call it a failure for the most part.
This is an interesting point, and I definitely wonder how it'll affect the tech industry in future.
> 4/ Subscriptions became the second scaled business model for web and mobile businesses, following advertising which emerged at scale in the previous decade.
Have they? They've done really well in some markets, but also done really poorly in others. People are definitely interested in them for music, films, TV, etc, and people + companies are interested in for them for certain products and services, but there are still many areas where they haven't done nearly as well as expected.
For instance, subscriptions still haven't worked in the media industry for news/journalism, as much as certain publications are hoping they would. They also haven't done too hot in the games industry either, if Google Stadia's performance is to be believed. Jury's still out for YouTube/Twitch creators too, with Patreon type services being the de facto monetisation method right now. Free with ads has usually beat out subscriptions when in competition with it.
> 7/ Technology inserted itself right in the middle of society this decade.
It sure did, though it's more like people took notice this decade, since the political tides went one way rather than the other.
> And the stagnation of earning power in the lower and middle class is absolutely the result of technology automation, a trend that will only accelerate in coming years.
It's effects on politics will accelerate too, as shown by this decade's move towards more extreme political parties and policies worldwide. This might not end well, especially with environmental issues caused by global warming/climate change at the same time.
A pattern I identified was that subscriptions work really well when property rights are strong and enforced strongly (TV shows, movies, music, books etc). It doesn't really work well when property rights can be bypassed easily or have a convenient alternative (you want to read NYT story without a subscription? read that story on some other outlet).
I think this reality will arrive hard and fast in the 2020's. The breadth and depth of regulation globally will depend on election outcomes. But it will happen either way. GDPR and CCPA are just a taste of what's to come. We'll see regulation in several key areas including:
- Social media's goal of increased engagement vs mental health.
- Cross border cybersecurity.
- More jurisdictions adopting GDPR/CCPA-like privacy laws.
- Cloud providers (consumer in particular) and vendor lock-in vs portability.
Some is already happening, but expect to see a ramp up. Don't interpret this as me being pro/anti regulation. Just a reality that's coming IMO.
I find it funny that there’re still so many CEOs that I speak to and they are constantly pushing the idea that tech should not be regulated.
Even if you’re right, there’s probably not a single politician out there who wouldn’t use that as a way to build their career on.
There is a lot of industry built around supplying the half a trillion a year military. They are well practiced in resisting reductions in funding. If the US becomes isolationist does that military might get turned inward.
Which option do you think is more likely
a) Stop buying so many toys
b) Overflow our sandpit into the rest of the yard
c) Play with our toys in other peoples' sandpits
Isolationism requires cutting out c as an option.