> Are you writing huge amounts of data?
No. All this data is reviewed by at least 1 human. It may be added by a crawler and skimmed for QA, but it is higher level data as opposed to event stream data. Small quantities of data currently in 10MB range but I suspect won't exceed 1GB
> How are you reading the data?
Easy to hold it all in memory.
> I'm assuming many of the columns will be empty for most records?
Yes. Currently 90% sparse. I expect eventually will follow the 80-20 rule.
> why you want to use ultra wide tables?
This is a long winded answer. The short version is "I don't know". It's a gut brain thing that I'm trying to move upstream to my symbolic brain and get the math onto paper. It feels to me like there is something "magic" in a wide dataset. I'm trying to formalize what that magic is, or maybe I'm hallucinating it.
Years ago I had the flash of insight (I wasn't the first, nor will I be the last) that all data structures, code included, could be represented with just words and indentation. I called my implementation Tree Notation and languages built on this syntax Tree Languages. (S-Expressions without parentheses is a 90% accurate description).
But the idea failed to catch on. It could be a bad idea. I'm still not sure. Perhaps it was missing something. I started collecting information about every notation and language I could find. This has helped me make Tree Languages better (though I still don't have a clear model as to whether the underlying notation is great or not), but also that dataset became a product itself and I shipped it at PLDB.com.
With the PLDB dataset, I feel like I can now quickly compute an accurate research cost estimate to any question about programming languages (and increasingly the answers are instant from the data already there).
I feel like there is a non-linear effect driven largely by having hundreds of columns, and I think it would scale where the product would be far better with another 2x of columns (and also fill in more missing data, of course). We are just kind of adding data and will have some empirical results as to whether that is right (perhaps it's wrong and the product won't substantially improve and the returns of adding more data will be diminishing—not sure). But I'm trying in life to take a more theoretical look at things to, and try to figure out formally why this gut feel is right or wrong.
One idea at formalization is that the value of a dataset increases exponentially with the number of independent columns (Value = n^2). Obviously that's the best case, the average case impact of a new column would be worse by some factor. There's probably some known math around this from the Decision Tree literature.
Another idea is that adding columns, even if sparse, is valuable because then you make known the unknowns. Wide tables could maximize the chances of collecting columns that offer a wide variety of perspectives to maximize the odds of finding the best models that most accurately predict the data.
Anyway, as you can tell it's not very clear to me yet, but I'm confident someone has written some math somewhere about the shape of the return curve from adding more columns to a dataset.