109 karma · joined November 17, 2014
https://docs.deeplake.ai/en/latest/deeplake.html?highlight=l... https://docs.deeplake.ai/en/latest/deeplake.html#deeplake.re...
Good observation, we've made sure that datasets can evolve better as you go. Specifically for your use case, you can query subsets of data and materialize it on the fly to be streamed, and then go back to a specific dataset "view" (i.e. saved query), as needed.
See how this works at around 5:50 here - https://youtu.be/SxsofpSIw3k
As for your last question, when the jpeg is appended using its file path, the compressed bytes get stored in the dataset without decompression/recompression. When the data is accessed as a numpy array, then the jpeg bytes are decompressed.
For researchers in Academia, our Growth plan is free. Since you work at a startup, the trial for Growth plan is for two weeks. If you want access, hit us up in the Community slack (slack.activeloop.ai - or you can just test the querying on public activeloop datasets!)
If you find any other points of confusion, please send them our way, and we will fix it, the community has been instrumental over the years in iterating on the product! :)
thanks a lot for the input, and thanks for trying us out. You know, it's always a work-in-progress, but we've actually done a major overhaul - this is our biggest release yet and I'd love it if you gave it a try, especially the querying feature we're super-proud of. :) https://docs.activeloop.ai/tutorials/querying-datasets
here's a couple of playbooks on how the new features (visualization + querying + version control) play together to solve complex workflows.
- https://docs.activeloop.ai/playbooks/training-with-lineage - https://docs.activeloop.ai/playbooks/evaluating-model-perfor... - https://docs.activeloop.ai/playbooks/training-reproducibilit...
Likewise, we curate a list of large open source datasets here -> https://datasets.activeloop.ai/docs/ml/, but our main thing isn't aggregating datasets (focus for HF datasets), but rather providing people with a way to manage their data efficiently. That being said, all of the 125+ public datasets we have are available in seconds with one line of code. :)
We haven't benchmarked against HF datasets in a while, but Deep Lake's dataloader is much, much faster in third-party benchmarks (see this https://arxiv.org/pdf/2209.13705 and here for an older version, that was much slower than what we have now, see this: https://pasteboard.co/la3DmCUR2iFb.png). HF under the hood uses Git-LFS (to the best of my knowledge) and is not opinionated on formats, so LAION just dumps Parquet files on their storage.
While your setup would work for a few TBs, scaling to PB would be tricky including maintaining your own infrastructure. And yep, as you said NAS/NFS would neither be able to handle the scale (especially writes with 1k workers). I am also slightly curious about your use of mmap files with image/video compressed data (as zero-copy won’t happen) unless you decompress inside the GPU ;), but would love to learn more from you! Re: pricing thanks for the feedback, storage is one component and customly priced for PB-scale workloads.
You can then visualize your datasets if their stored on our cloud, in AWS/GCP, or you can drag and drop your local dataset in Deep Lake format into our UI (https://docs.activeloop.ai/dataset-visualization)
We do, with version control, Python based dataloader and dataset format being open source! Please check out https://github.com/activeloopai/deeplake.
From what we are seeing in the market, both domains grow, but with an overlap, and it's expanding, too. I think while BI/Analytics would still be a major space, we would see more DL-based novel applications generating increasingly more business value (i.e. self-driving cars, robotics, agritech). After all, even in VERY traditional workflows/companies like economic growth estimation, we're seeing DL being applied (e.g. they look at nightlight satellite imagery to estimate economic growth/urbanization).
So to answer your question, for some parts, I think it would be the former (complement), and other applications it would call for replacement (particularly in the cases where companies use multi-modal data).
Fair point regarding the unlabeled/unstructured data. One could also argue that labeled data isn't going to be a prerequisite forever (see https://ai.facebook.com/blog/the-first-high-performance-self...). We see a very sharp rise in unstructured data use for ML (especially a large spike caused by large language models like Dall-E 2 and Stable Diffusion). In my opinion, the majority of the novel use cases are outside of big tech, and we also see a trend in "legacy" companies like media, manufacturing, etc. start building dedicated ML teams. The industry is still nascent, but it is growing fast. Frankly, we see the pain points we're solving resonate with so many more companies than just a year ago.
Agree re Snowflake/Databricks, they are partners rather than competitors. We sit on top of S3/GCS or other blob storages and currently are competing with various in-house solutions that ML scientists built themselves. I do see your point regarding large foundational models that would be only fine-tuned on the tail end for various use cases. I believe there still would be still companies building foundational models from scratch (currently at 5 billion images) so they can serve more application-specific products and unstructured data generators that partner with those companies creating a good enough market for the tool.
Our main competitive advantage against the players you've mentioned is just that - our bet is that Deep Learning will overtake traditional BI workflows (especially with >90% of data generated today being unstructured), and we've been preparing for it. Traditional "BI datalakes" are pretty inefficient when it comes to storing the data specifically for deep learning workflows. They currently also lack an entire suite of key features (visualization for those data types, query engine based on tensors, etc.) to be able to successfully convince the potential users.
As a matter of fact, we're seeing not only adoption from AI-first companies/startups who are building their infrastructure from the ground up, but mature companies who are hitting the limits of the traditional setups.
Keeping that in mind, we're working on making the onboarding for such companies much easier, so their cost of switching to a more efficient/performant setup is much lower.
As for Databricks specifically, we see them more as a complement, rather than a competitor.
You are right my claim that y'<y is slightly weak (was based on "gap is not negligible" assumption, see below).
"gap is not negligible" - means if y' gets near to y, then y will get even higher and there will be always a market gap, which I think you disagree with.
Based on your suggestion, on extreme scenario I would soften my claim to y'=<y, without us making profit. :)
If the price goes down, customers will be able to set lower price. As long as GPU holders profit margin is high enough given electricity and maintenance costs, they will do the compute.
If the marketplace matures, pricing of mining, deep learning, rendering and other tasks will be driven by the market. At Snark AI we are working on towards creating this marketplace that will provide optimal benefits to all parties.
If you want to deploy large-scale computation and significantly reduce your costs, we can help by running mining at the same time under your consent. This only applies to Deep Learning inference.
Regarding WebGL, actually that is an interesting point, would like to know more about the use case.