A Critical Field Guide for Working with Machine Learning Datasets
knowingmachines.org
knowingmachines.org
100 years of statistical advances have not been able to solve image recognition at scale until convolutional neural networks. Large language models may be somewhat Bayesian but they do not resemble any traditional Bayesian techniques except at the highest level. Clean data and good sampling are incompatible with the real world outside of laboratory conditions (or large budgets).
[1] https://ai.googleblog.com/2021/06/data-cascades-in-machine-l...
It is still taught at university if you do a master degree or PhD at a decent place.
But I agree that many juniors have no clue. Especially those who are self taught and join in from a different field
Anyone have a suggestion for storing large amounts of ML data, training sets or otherwise? I’ve been using FeatureBase and weaviate, and would be interested in learning about other solutions.
Would love to get some feedback on it!
Though I think one common approach is to just dump most data in an S3-compatible datastore of your choice (there's seaweedfs or ceph, or the cloud provider of your choice). Specialized databases make a lot of sense when you need features like vector similarity search though.
Milvus (https://milvus.io) and Feast (https://feast.dev/) are two of the most well known vector databases and feature stores, respectively.
The design gave me a bit of anxiety, it was so convincing I was worried about clicking into one of the cells!
Anxiety included - I also love the layout.