Read “Data and Reality”
buttondown.email
buttondown.email
Right now, I am wrestling with exactly the same issues: designing data structures for representing real-world entities, discovering the roles they play (in a financial decision in my case), gathering data about the entities (depending on their roles), confirming the relationships between them, and validating attributes about them, based on primary data with differing degrees of reliability.
I've been through various iterations, painfully getting better at representing "what is true and what don't we know yet".
Reading "Data and Reality" just now was like a distillation of my endeavour, all in one place. Plus lots of things I haven't properly considered yet.
I can see why many people find it dull: if you haven't struggled with representing these distinctions yourself, it can well sounds like pointless nitpicking and overcomplicating straightforward situations.
So a big thank you to the original poster, Tomte.
In particular, it gives a reference to Bill's web site http://www.bkent.net, which includes a list of his personal and academic [1] writing over 35 years.
I was actually astonished to find his web site still running, 17 years after his death in 2005. How many of our own digital artefacts will have such longevity?
I just dug it out again. The start of Chapter 11, over 190 pages into it, may be hinting at the issue:
"Thus far we have been largely critical, and negative. We have identified problems without really suggesting solutions. Can we identify an appropriate set of elementary concepts that will on the one hand serve as a general base for modeling information (in our limited use of that term), and on the other hand be an appropriate base for computerized implementations? Let us try."
Since it is stressed that the book is philosophical, I have to say that the most paradigm-shifting ideas about modeling reality I've encountered so far come out of second-order cybernetics and especially the work of Ranulph Glanville. His "Black Boox" series is hard to find online but highly recommended.
Would you mind expanding a bit more on how second-order cybernetics was paradigm-shifting to you?
Could you possibly give a brief description of what it's about?
I hate these kinds of arguments. It assumes that every book has an equally likely chance of being good. Books being popular and books being good generally have a correlation.
Most popular books become obscure books, given enough time. (Reading Les Miserables, I had to consult the footnotes repeatedly for Victor Hugo's allusions to then-popular novels and novelists, almost none of which remain popular today.) If you're a serious reader the chances are diminished that your favorite book is something that has just recently been written. But it's also likelier than random chance that your favorite book was once popular, because as you note there's generally a positive correlation between quality and popularity.
No it doesn't, it just has to assume that
P(good|obscure) > - P(popular)/(P(popular)-1)
Or, more practically, when P(good|obscure) is just a hair more than P(popular)
Let's say all popular books are good, so
P(good|popular) = 1.
Then we'll say 1/1000 books are popular.
This means P(popular|good) == P(obscure|good) (i.e this is when the number of good books that are popular equals the number of obscure books that are popular.) when P(good|obscure) = 1/999. This true if we assume that P(good|popular) = 1, which is the highest value it can take. If this number is lower than this constraint is reduced, so we can take this as an upper bound of the relationship.
So knowing nothing about the rate of goodness among popular books, we can assume as that there is a huge number of obscure books, and a book is just reasonably more like to be obscure and good, than it is to be popular, the we can confirm that there are more good obscure books.
This constraint is much less demanding than assuming all books are equally likely of being good.
The math is the same as hot/crazy/marry plots: draw popularity vs. goodness (random correlated dots: more popular more goodness). Define which books may be favorite e.g., above G line goodness a book has a chance to become favorite, define threshold for obscurity e.g., less than P popularity level. Consider what happens to the number of obscure books the more you read the more of them can become popular.
If you model it with a single threshold then the math is the same as for «the better at programming competitions the worse at “some other metric for coders” among google hires» (imagine you sum two metrics and hire only those who have the sum above a threshold -> you get the positive traits in reverse relationship).
I haven't read the third updated edition, with changes by a different author. I've heard mixed reviews for that one. I'd be interested in any opinions on it from anyone whose read it and a prior edition.
In summary, it's not great.
Author’s full name is William Kent (updated by Steve Hoberman in 2012 ed).
Title is: Data & Reality …(2012)
The original was published in 1978 and is very expensive (on Abebooks) at $242 minimum, but the new version reasonably priced.
Sorry, I don’t see your guidance.
What is date on the 2nd edition that you recommend? I have scoured Abebooks and not found a single copy. And generic search failed.
There is an unstated dubious assumption that popularity is independent from goodness here. Not that I disagree with the conclusion.