Graph-Powered Machine Learning at Google
research.googleblog.com
research.googleblog.com
Big data and in this case the relationship (graph) between big data points are whats needed to make great ML/AI products. By nature the only companies that will ever have access to this data in a meaningful way are going to be larger companies: Google, Amazon, Apple, ect. Because of this I worry that small upstarts may never be able to compete on these type of products in a viable way as the requirements to build these features are so easily defensible by the larger incumbents.
I hope this is not the case but I'm getting less and less optimistic when I see articles like this.
The one thing that makes me feel slightly better about this is the recently announced AmaGoogFaceSoft partnership on training data open sourcing [2].
I say slightly because they still control the pipes and what is released - so they'll never likely release the good data, also cause it would probably cause privacy issues if they did. It's incumbent on us as the little guys and girls to make new data pipes through innovative products.
[1] https://medium.com/@andrewkemendo/the-ai-revolution-will-be-...
You are right that Google has a lot more raw data, but that is much cheaper to obtain, and depending on your problem domain, there could be a bunch of public datasets that are available for you to cobble together.
Some of these big tech companies don't necessarily have the right training data for their applications either. For example, Adam Coates from Baidu outlined some of the ways that they generated a better training set for their speech recognition use case here: https://youtu.be/g-sndkf7mCs?t=3817
Errr, not really? I mean, at a minimum you're making a somewhat controversial statement - lots of people (me included) wouldn't automatically agree that the rich get richer, especially if you're thinking "only" the rich get richer, which is what is implied by the comment above. (As evidence, see established companies that go bankrupt, and new startups which rise from nothing to become dominant players, much like Google itself, which has existed for a relatively brief period, or something like Uber).
So let's hope that data capitalism is like regular capitalism!
But the dynamic I have described is certainly there. And it can become dominant.
http://www.latimes.com/business/hiltzik/la-fi-hiltzik-ft-gra...
On growing income inequality and the transfer of wealth to the 1%.
You can argue if it's right or wrong but it's a bit of whopper to say it doesn't happen.
As a teaser, one important thing is to not share information (i.e. Don't blog about financial "hacks"). Call it a corrupt moral compass or whatever you like, but the fact is this behavior is rampant.
That something so important is decided at an early level of one's life that has wide reaching implications is quite frankly a disgrace.
Only the most endeavours break the cycle and they are too few and far between that are not really encouraged on most levels of society around the world due to the norms.
the vast majority of financial tricks require scale (of capital) - e.g. itemized deductions, business expenses, even tax-loss harvesting - but they aren't secret!
Also, I don't think anyone knows the exact causes of this growing inequality - some libertarians, for example, would argue that it's excessive regulation that is causing part of the problem - which isn't capitalism, it's the opposite of capitalism.
But most importantly, note that I specifically talked most strongly against the idea that only the rich get richer. Just as an example, how many of today's billionaires are relatively new to the game? It's not all of them, not by a long shot. And while some minimum level of "being rich" is certainly a factor (in the send that everyone who lives in a 1st-world country is rich), there are lots of people on that list who started with relatively little.
To truly see who is more fit, I propose that we take a group of 12, 6 from each class, and drop them off on an uninhabited island. We will places weapons and traps at strategic locations. We shall call this experiment Project Craving Romp.
One of Piketty's points is the relationship between the rate of growth of the economy, `g`, and the rate of return on capital, `r`. If the rate of return of capital exceeds the rate of growth of the economy, that is `r > g`, then inequality of wealth increases over time. As they say, "the rich get richer".
These conditions (`r>g`) have been observed from historical data under normal conditions (excluding events like world wars), and we expect them to continue into the future, particularly with projections of declining or stabilising global economic growth rates [1].
The period after the two world wars (with relatively high equality) is unusual in history and over we're seeing inequality of wealth increase, inherited capital becoming an increasingly relevant factor, etc.
[1] let's ignore the other perspective of whether continual economic growth is actually feasible or desirable given the reality of hard environmental limits.
Wealth is effectively "liquid power" and "power" is the ability to get other people do to what you want. One of the absolute most obvious things you'd want them to do is... give you more power.
So I think it's a pretty natural outcome that, absence other forces, any power imbalance will tend to magnify over time.
Times in history where the entire economy is growing very fast basically mean new power is raining down on all people uniformly. That will tend to reduce disparity in the same way that adding the same positive number to both the numerator and denominator leads to a fraction closer to one.
Of course, this simplified model treats every person as an island. Where the story gets more complex is when you consider people working together in a group. And I think through most of history when you've seen power imbalances get reduced, it's because you've seen people work together to form groups that have greater power than the smaller number of individuals they are pushing against.
One of the things that really scares me about the US today is how much we've culturally lost that ability to organize and work together. And, of course, the small number of increasingly powerful people and groups like it that way, as they always have.
- https://research.googleblog.com/2016/09/announcing-youtube-8...
- https://techcrunch.com/2016/10/05/udacity-open-sources-an-ad...
There are also more paid, private dataset services now:
- https://fantasydata.com/pricing/nfl-data-api.aspx
I could see there being a lot of opportunity for startups that democratize data, which would then unlock more startups to build novel ML architectures on top of.
Google has thousands of real robotic manipulator arms to train their reinforcement learning algorithms. Algorithms trained using the OpenAI gym are practically guaranteed to fail in the real world since OpenAI gym is (designed as) a perfect simulator. Don't even get me started on internal image and speech recognition datasets collected from YouTube, search indexing, and Hangouts/Duo calls.
There is no way in hell you can compete with this. The closest you have is Facebook. But they've built an AI that has just learned how to play Go. Let alone defeat a world champion. Hell, Google is already working on beating humans at Starcraft.
Wow, what's with all the fatalism around here? C'mon, people have been making variations of this argument for hundreds of years, but small companies still keep coming along and knocking the KOTM off its perch. There's always a new way to do things, a new angle, a new hack, something that hasn't been done before, that can (somewhat) level the playing field (for a while).
2) Write a web scraper and run it for a day on an AWS instance.
3) Download any of the publicly released data sets. There are lot of interesting ones out there, including medical imaging, etc.
4) Hack a video game engine to generate data.
5) Write your own data generator (could even be a generative NN model). Not all data can be simulated, but when it can be, doing so is incredibly powerful.
A few of the above options were not available just 5-7 years ago. Data is becoming more publicly accessible, not less.
Here's why not to be sad: Retrofitting and its variants are quite effective, and surprisingly, they're not really that computationally intensive. I use an extension of retrofitting to build ConceptNet Numberbatch [1], which is built from open data and is the best-performing semantic model on many word-relatedness benchmarks.
The entirety of ConceptNet, including Numberbatch, builds in 5 hours on a desktop computer.
Big companies have resources, but they also have inertia. Some problems are solved by throwing all the resources of Google at the problem, but that doesn't mean it's the only way to solve the problem.
Short of signing up to the mailing list on Google Groups is there any info out there on the net that tells me:
a) How ConceptNet relates to efforts like Cyc's OpenCyc and the big G's Knowledge Graph?
b) How to add domain specific concepts, attributes, and relations to ConceptNet
Both these things are the first that pop into my mind but the website: http://conceptnet5.media.mit.edu/ doesn't seem to cover them.
In relation to OpenCyc: ConceptNet doesn't have CycL or any sort of logical predicates -- the assumption is that you are just doing fuzzy machine learning over ConceptNet. I'm pretty sure you need the rest of Cyc to do anything interesting with CycL, though.
ConceptNet takes advantage of linked open data when it can, so it does contain attempted translations of the facts in OpenCyc.
Compared to the Google Knowledge Graph: well, you can't really use the Knowledge Graph without being at Google, or making a research agreement with them, right? But from what I've seen of it, and of its predecessor Freebase, I think it focuses a lot on named entities: particularly things you can look up on Wikipedia or things you can buy. Which is fine information. I'd say it's a different segment of linked data than "what words mean", which is ConceptNet's focus. So maybe think of it as more like a bigger WordNet than a smaller Knowledge Graph.
How to add domain-specific concepts to ConceptNet: the unsatisfying answer is that you can get the code for ConceptNet and alter the build process to include new data sources (try it on the 5.5 branch if you attempt this). And the startuppy sellout answer is that doing this automatically is one of the things my company Luminoso is for (http://www.luminoso.com).
Thanks for asking, and I'll try to clarify this kind of stuff when I deploy the new site.
But I do think this is much more like retrofitting than like label propagation. It's the vectors that are being propagated, as I understand it, not labels.
This used to be the case. But companies nowadays are often more structured like clusters of small startups, especially for innovative products and moonshots and such.
This is a step in the right direction, and from experience, a small research-only startup could achieve very similar results.
More than likely, the amount of data is a detriment to creating a better system.
I agree right now big data centralisation is a problem indeed, but I don't believe it will last. I mean, those big companies will remain but Blockchain tech is fairly likely to take a big spot in there with them, I believe.
In the retrofitting paper cited in the comments there is a process of smoothing, that is feeding back the message or information to update the states of the graph (example in the modern ai book). It doesn't seem anything new.