Is that typical for this kind of work?
Is that typical for this kind of work?
Anyway, the diversity in customer schema leaks out into the Hadoop schema, where we'd much prefer to give customers data using column names they're familiar with, and we also want to give them rows from all their different schemas in a single table (because many schemas have overlap by design). The superset of all schema columns is large, however. The problem can be overcome with more tooling - defining friendly views with explicit column choice - but having the option to implement that (and go to market sooner), vs a requirement to implement that, adds up to a distinct advantage for tech that can support the extra columns.
(I know, in a column store having related data in another column isn't actually close together; but it can be stepped through at the same time, it doesn't need a join to be correlated, it's correlated naturally.)
I understand though that sometimes customers give you a rotting dumpster of data and ask for critical insight into their operations.
That's why I always find it hilarious when HN goes on about just using PostgreSQL or some other SQL database for everything when they don't understand the use case. They simply doesn't work in these scenarios.
Why so many columns? I would expect there would be some way to break the data down or transpose it somehow to make it more manageable. At that point are you even really dealing with a database as most people use the term?
It has to be one table because you need to get attributes of a customer very quickly (single digit milliseconds) in order to respond with the next best action e.g. show this advertisement or route them to this call centre person.
And of course this is a database. It's all of the information about a customer in one place.