Machine Learning, Kolmogorov Complexity, and Squishy Bunnies (2019)
theorangeduck.com
theorangeduck.com
Shannon information depends on population statistics of messages, usually has a 'shallow' and domain-specific interpreter (to turn encodings back into messages, e.g. like a Huffman tree), and tries to minimise the average length of a set of encoded messages. In this case the NN is the interpreter, the length of the encoding is the number of deformation axes, and the population is the training data (which we hope is representative of the real data).
Algorithmic information is different. It depends on patterns within an individual message, rather than across a population. The interpreter is a Universal Turing Machine (or equivalent), which is general-purpose. Encodings are programs which output a message, and can be arbitrarily 'deep' by running for an arbitrary amount of time.
Another way to look at it is: using PCA to solve the problem of 'automatically decide which deformation axes to include' is slightly interesting; much more interesting is the problem of 'how can we automatically decide that this data should be modelled as a set of deformations along various axes?'. That's a much harder problem, and tends to require a lot of search through program-space (whether NNs or GOFAI).
In my experience, general practice w.r.t. Shannon information is to manually pick a model with some numerical parameters (Huffman tree, deformation-axes-NN, etc.), then set those parameters to work well on an example population.
General practice w.r.t. algorithmic information seems to involve searching program space for a good model. The adjustable part is the language, which can certainly be manually tailored to the application, but doesn't tend to be the bulk of the work (e.g. automatically searching for an efficient language tends to be redundant once we're doing program search; since those programs will contain their own abstractions in any case).
PCA was my introduction to data science and it remains one of my favorite tools to pull out. I'm really surprised by the number of data scientists that I meet who never use it, or have even never heard of it. Then again, even Data Science From Scratch dedicates like two pages to the subject, so maybe it is not something modern data science programs spend any time on.
In fact, SVMs literally make that tradeoff by using the Kernel trick to save up on the computational costs of projecting to higher, possibly infinite dimensions. As we know, projecting up to higher dimensions can provide us with a 100% accuracy on the training data, but of course, that's overfitting and we use constrains/support vectors to regularize the learner.
Did the PS5 have a chip that could execute small neural networks? My PC doesn't have one right now, but if I bought the latest in GPU there is a dedicated chip already. This is all very new to me.