Show HN: TheBigDB - A simpler open database of facts
thebigdb.com
thebigdb.com
First of all, let me say that I'm glad more people are thinking/working in the space of triples. Even unstructured ones like this.
But when there's no semi-strict schema, it gets really, really tricky. Free text is hard, and actual meaning is hard to separate. (I say semi-strict, as Freebase is schema-last -- feel free to create your own! -- but has some level of enforcement)
For specific domains you may be okay with tags. And for some limited applications it probably works great. Triples are cool!
But when you start talking about larger, broader, datasets, ones that no one person or small group can curate, you're going to start running into collision.
There's certainly an argument to be made for metaschema -- https://developers.google.com/freebase/v1/search-metaschema -- and crowdsourcing these sorts of things could be interesting.
I think there's a lot of interesting work to be done. But I doubt that this is "better" per se, or at the very least, is little more than a toy.
And hey, I built such a toy graph engine once upon a time (be gentle -- it was really a demo hack) https://github.com/barakmich/jgd -- you can even query it with Freebase's old MQL. (Which I have mixed feelings about, but is cool in its own way)
I guess my argument is, don't throw the baby out with the bathwater. And feel free to ping me for more!
Second: It is not intended to be a Freebase killer, but it is certainly better for me and what I needed it for: building denisthebot.com - which can answer questions about a whole lot of subjects while being completely agnostic about what "categories" these subjects fit in.
Third: Since we all know no-one is going to change anyone's mind here, I won't discuss the merits of a DB structured the way I built it. But calling it a toy is quite flattering, I know some very powerful stuff that were called "toy" for some time :) Also, it would be a very easy toy to use, which will suit a lot of people, according to the emails I received.
Also, you're right: "don't throw the baby out with the bathwater".
> Simpler structure: There are no datatypes, namespaces, lists, domains. Just ordered nodes. Having a dead simple structure like that allows developers to quickly and intuitively know how to access the info they want.
I don't see how this makes it simpler or intuitive at ALL. If there's no convention as to whether I should search for "born on" or "born_date" or "year_born", or whether the date will be "1900-08-01" or "08-01-1900" or "1900/08"... then how is this supposed to be useful?
The central problem is, there are lots of textual ways of describing the same thing. Without standardized datatypes and standardized tags, it quickly becomes a messy, useless free-for-all.
I don't see how TheBigDB gets around this. The FAQ explains how it's different from Freebase/Wikidata, but I don't at all understand how it's supposed to be better, or even as good.
Right on the front page they have their biggest flaw in the example:
Apple weight 150g
Ok, so what about the company Apple? How are they semantically distinguishing one item from another? I guess Steve Jobs is Founder of a Fruit that weighs 150 grams.
It will be released when it will be the right time for it. Thanks for your comment!
You'll then need to start assigning contexts to entities, at which point you're right about where Freebase begins, in principle. Even if we assume you use Wikipedia-esque strings like "Apple (company)"
Triples are easy. Semantics are hard.
I don't want to talk too much about a structure that I haven't released yet, but let me assure you there's a way that doesn't end up where Freebase is. (Not that there's anything wrong with that.</seinfeld>)
And yes, it is different, "better" depends on what you're trying to do with it :)
Maybe I'm just impatient, and something mindblowingly awesome, based on Cyc, is around the corner.
>Lenat was frustrated by Automated Mathematician's constraint to a single domain and so developed Eurisko; his frustration with the effort of encoding domain knowledge for Eurisko led to Lenat's subsequent (and, as of 2008, continuing) development of Cyc. Lenat envisions ultimately coupling the Cyc knowledgebase with the Eurisko discovery engine.
I don't know what he intends to do with it from there, but it could potentially make for some very powerful AIs.
The first one because it makes people actually want to use the service, the second one because it helps having up to date data about all sort of things.
Which date is "10/3/5"? Is it March 5th 2010? March 5th 1910? March 10th 1905? March 10th 2005? October 3rd 1905? October 3rd 2005? (or another century entirely, though the 20th and 21st would be most likely). And don't think the "/" vs. "-" as separate is sufficient to tell them apart.
Aand you'll find a lot of other variations - I'm used to writing 10/3-5 for example... But I'm not even consistent, I might write 10/3/5 or 10-3-5, or 5/3/10 / 5-3-10; anywhere I want to be explicit, I would write 2005-03-10 exactly because I'm used to seeing so many ambiguous dates that can't easily be resolved.
What about the value 5.123? Is it a floating point value with "123" after the decimal point, or the integer 5123? The "decimal point" is "," in many countries, and the thousand separator is usually, but not always, "." in countries that use "," as the decimal marker. If you treat things as "just text" you are going to have to potentially deal with dozens of different combinations of decimal points and quantity markers (depending on country, the markers don't all occur only every 3 digits to the left from the decimal marker...)
Interpreting small text fragments is fraught with a near infinite number of obnoxious details like this, and part of the problem is that even few people know most of them and will be unable to quickly resolve ambiguities without cross referencing with other data (or worse: they think they know, or don't even recognize that there is an ambiguity in the first place)
I think that's far better than voting. Voting for facts amounts to relying on a logical fallacy: appeal to the majority. [2] (Voting is fine for popularity contests, or things that can only be matters of opinion, but facts?)
If no, then sorry, freebase is vastly superior IMHO - from user's point of view I don't see a point in a crowdsourced proprietary database (even if API is currently free).
E.g. they explicitly say on their website: "Copyright does not protect the facts or ideas underlying the creative expression. So, Creative Commons licenses do not apply to ideas, factual information or other non-creative elements that are not protected by copyright."
Courts in many jurisdiction explicitly refuse to accept "sweat of the brow" arguments for copyright, explicitly requiring an element of creativity. Some countries do have "database copyrights" that can protect arrangements of facts in certain ways, while in others a collection of straight up facts can not be copyrighted pretty much no matter what.
So, if I understand correctly, it let people crowdsource any kind of structured and descriptive data ?
For example, as you can see on http://denisthebot.com, having that kind of API makes it trivial to built smart question-answering engines.
Can I send how many requests I want?, I think you might mean Can I send as many requests as I want? ?
Facts are verifiable statements and things that are verifiable are verifiable repeatably. This is not a verifiable statement.
Yes, of course.
> So, no statements about political geography at all, then.
Absolutely not, you just have to make time explicit.
Edit: It seems we're arguing about what a priori knowledge is capable of serving as a base for factual deductions. The Kantian approach is to say that we all agree on time and space and everything can be based off of these self-evident truths. I think there is not such a clear boundary between objective truth and induction.
Edit2: I'd also like to take this moment to point out that "you're" is the proper contraction of "you are", since we're getting all semantic.
> I'd also like to take this moment to point out that "you're" is the proper contraction of "you are", since we're getting all semantic.
I know, it annoys me too. By the time I'd realized it was too late to edit. Typos happen.
But they definitely were true. I thought you were making a distinction between 'something that is true' and 'something that is a fact (ie: is unchangingly true)' which I don't think most people make.
The easiest example is this: We're in 2010. X has been married to Y since 2009; In the DB it would be represented exactly like that: a "from" time period, no "to" time period, meaning "it is still true now".
They divorce in 2013. X married to Y from 2009 to "now" isn't true. It should be downvoted. but X married to Y from "2009" to "2013" is actually true. It should be created and upvoted. That fact surely won't change overtime.
This is not a fact database. This is:
X Married Y 2009-01-01
X Divorced Y 2013-01-01
This way you can represent any number of facts: X Married Y 2013-02-01
Think how convoluted it would be for you to represent X and Y marrying again in your example, if not outright wrong (because you update/delete information).This is akin to your bank storing the total amount of your account in their database, instead of the transaction history and deriving the total from that.
There's nothing wrong with that, until there is.
Consider this situation: X marries Y, divorces, marries again. Now you would have the date of divorce prior to the date of marriage. How do you make sense of this data?
> And since the DB handle time periods, with my way you can actually search who X is married to with one request, without checking if he divorced, if Y is dead, disappeared, or else.
The fact Y is dead doesn't mean X wasn't married to it. So despite the obvious technical implications of keeping all this data in sync (e.g., Y dies so you would update X to reflect that?), you're simply obliterating information. A fact is immutable, therefore a fact database, by definition, only appends, never updates.
It is, indeed, a hard problem.
I like the new approach though. You appear to be focused on simplicity and ease of use, and hopefully your find ways to fix the resultant problems. For example, some fancy graph theory might be able to determine that the graph node "Apple" refers to two different ideas.
Check, for example, Wikia, which contains Wikipedia for TV Shows which are more complete than Wikipedia.
If you're worried about competition, release under AGPL.
Similarly for the data, to the extent restricted by copyright and such, release under CC0 (same as Wikidata). Or if you want to be onerous, ODbL.
Some neat ideas in the project in any case, good luck!
As you said, conventions help mitigate the problem a little bit but the end user can hardly be expected to stick to best practices.
I have hope though. This is a problem worth solving.
As others have pointed out, some kind of conventions must be established around the semantics, and something must be done to avoid redundancy (which leads to inconsistency) and ambiguity.
I agree with those criticisms, but if the community also helps develop the schema, it will be interesting to see. What collisions will happen? What will be the result of queries that reach far across disciplines?
A suggestion: it needs a demo query box on the site. Shouldn't be too hard to let a rate limited IP address throw a few keywords at it and spit back results. I'd like to see what the db contains before I invest too much time (how many topics, how many facts, etc).
In the meantime, you can just look through http://browser.thebigdb.com what the DB contains, or just "gem install thebigdb" and start copy/paste the code examples to see how the API really behaves.
Thanks for your suggestion!
(Voting democracy may help prevent people from being oppressed in certain ways, but it isn't much of a truth-discovery mechanism.)