That's 100,000 english wikipedias or 50 googles.
That's 100,000 english wikipedias or 50 googles.
You can retrieve the exact information even faster. From an (x,y,t) coordinate, you can find the exact index in constant time
> What sort of dataset are you indexing that has trillion entries?
It doesn't say
> What sort of dataset has a trillion entries?
Not that this is a practical example because we technically cannot index all cells in each body. But even if such an algorithm being studied today might be useful one day when we do capability to collect such data
If I was building it from my 5 minutes of googling, using 15TB nvme u2 drives, and easily available server chasis, I can get 24 drives per 2u of a rack. That's 360 TB + a couple server nodes. So ~6u per PB. A full height rack is 42u, so 6-7PB per rack once you take up some of the space with networking, etc. So dozens is doable in a short datacenter row.
Realistically you could fit a lot more storage per U, depending on how much compute you need per unit of data. The example above assumes all the disks are at the front of the server only, if you mount them internally also, you can fit a lot more. (see Backblaze's storage pods for how they did it with spinning disks).
Dozens of PB is not that much data in 2023.
Yes it is. Just transferring it at data center speeds will take days if not weeks.
As far as I am aware Google doesn't publish any statistics about the size of its index, which no doubt varies.
for each photo, extract and index SIFT vectors
https://en.m.wikipedia.org/wiki/Scale-invariant_feature_tran...
i used it as an example of a way a dataset might have a huge number of feature vectors.
curious if there are better ways to do this now.
I'm not in the we scraping/search area, so idk about the 100B website thing (other than it's Google order of size), but the encoding would take some mega amount of time depending on how it's done, hence the suggested sum of GloVe chunks (potentially doable with decent hardware in months) rather than throwing LLM on there (would take literal centuries to process).
Think of it like an accountant. Weights are all of their experience. Prompt is the form in front of them. A vector database makes it easier to find the appropriate tax law and have that open (in the prompt) as well.
This is useful for people as well, like literally this example. But the LLM + vector combination is looking really powerful because of the tight loops.