Good observation, we've made sure that datasets can evolve better as you go. Specifically for your use case, you can query subsets of data and materialize it on the fly to be streamed, and then go back to a specific dataset "view" (i.e. saved query), as needed.
See how this works at around 5:50 here - https://youtu.be/SxsofpSIw3k
As for your last question, when the jpeg is appended using its file path, the compressed bytes get stored in the dataset without decompression/recompression. When the data is accessed as a numpy array, then the jpeg bytes are decompressed.
For researchers in Academia, our Growth plan is free. Since you work at a startup, the trial for Growth plan is for two weeks. If you want access, hit us up in the Community slack (slack.activeloop.ai - or you can just test the querying on public activeloop datasets!)