Some questions/observations:
- Python API is great and definitely more useful than CLI when it comes to regular use. But CLI would also be nice for some operations, such as import/export, see history metadata, etc.
- Just a nit, but if I didn't already by know about activeloop I might be less interested due to the "deep lake" branding. There's already so much "data lake" stuff that I have no interest in since it's historically not useful for image data.
- In academic contexts, it's typical to have a static dataset and always use that. In my use cases, typically I have an every-growing "raw" dataset and extract/preprocess subsets periodically for (re)training, annotation, etc. I typically consider these subsets ephemeral, as I can regenerate on demand. Last time I tried activeloop, it seemed like it was more geared towards the static dataset use case, but browsing the docs now it seems like there's more consideration of the latter case, so I'll have to look at that.
- In the examples I've seen, it seems like if a JPEG image is added to the dataset with jpeg compression, it first gets decompressed, so it would have to be recompressed - a lossy operation, right?
(edit: formatting)