However, it concisely represents a manifold in a much larger dimensional space and effectively captures most of the information in it.
It may be (and is) lossy, but don't underestimate the expressive power of a deep neural network.
However, it concisely represents a manifold in a much larger dimensional space and effectively captures most of the information in it.
It may be (and is) lossy, but don't underestimate the expressive power of a deep neural network.
It's dimensionality reduction. You cannot recover the original object. It's like using a shadow to reconstruct the face of the person casting the shadow.
Note this has nothing to do with the expressive power of a deep neural network. You are by definition trying to throw away noisy aspects of the data and generalize a lower dimensional manifold from a high dimensional space. If it's not lossy, it won't generalize.
[Edit: and that the salient characteristics are likely contained in the model.]
There is a real issue here of whether or not they should be allowed to keep a model trained from ill-gotten data. But the way I would think about it is: If you steal a million dollars and invest it in the stock market, and make a 10% return, what happens to that 10% return if you then return the original million? That's a much better analogy for what's going on here. They stole an asset, and made something from it, and it's unclear who owns that thing or what to do with it.
Regardless, I still think having the most relevant features already extracted is all they need to ask many of the questions they might want to. The point is that that’s still quite bad.
Makes me think of the Simulacrum[1]. "The map is not the territory."[2]
1. https://en.wikipedia.org/wiki/Simulacra_and_Simulation
2. https://en.wikipedia.org/wiki/Map%E2%80%93territory_relation
SSN is a lookup key into the raw data. Dimensionality reduction is by definition lossy since it's used in scenarios where: rows of data = n <<< m = number of features
And, for example, where someone's proclivity on the exploration/exploitation spectrum, if you will, (IE, how strongly do they respond to fear-based messaging) falls is probably quite predictable from a spectrum of likes.
Cat pictures may be less informative, but not all of these people clicked exclusively on feline fuzzy photos.
Is this falsifiable? It reads like a tautology to me.
On this tangent, IP ownership for deep learning models is interesting - how to you prove (in court) someone has/hasn't copied model/stolen a training set? If you fed someone else's training/model into your system, how easy is it to prove? Will we see the equivalent of map 'trap streets' in trained CNN models?
Which led me to: https://medium.com/@dtunkelang/the-end-of-intellectual-prope...