From what I understand Logica compiles to SQL so it can run on BigQuery. I don't
think that's putting it "on top of SQL".
I think it's a bit confusing that datalog is always discussed in the context of
databases and as a "query language" etc. In fact it's a subset of Prolog, so it
really belongs to the subject of logic programming. It doesn't help that Prolog
programs themselves are implemented as databases and that Prolog programming
uses terms such as "query" that blur the waters about exactly what one is doing.
I confess I don't have a background in databases and so I only understand the
very basics about SQL's semantics, which is the Relational Calculus, but as far
as I understand it, RC is a subset of predicate logic (a.k.a. first-order
logic). Prolog is itself a different subset of predicate logic, Horn clause
logic; and Datalog is a subset of Prolog and equivalent to SQL in expressive power.
Very briefly, every expression in Prolog is a Horn clause. A clause is a
disjunction of literals. A literal is an atom, or the negation of an atom. An
atom is an atomic formula, a predicate symbol followed by a number of terms in
parentheses where the number is the "arity" of the predicate. Terms are variables, functions or constants.
For example, father(john, bob) is an atom of the predicate father/2, where
"father" is the symbol and "2" is the arity.
An example of a clause is grandfather(x,y) ∨ ¬father(x,z) ∨ ¬parent(z,y). This
is a disjunction of one positive literal, grandfather(x,y) and two negative
literals, ¬father(x,z) and ¬parent(z,y). By the rules of logical connectives,
the same disjunction can be written as an implication: father(x,z) ∧ parent(z,y)
→ grandfather(x,y). By Prolog convention also observed in Datalog, implications
are written with the positive literal first: grandfather(x,y)← father(x,z),
parent(z,y). The left-facing implication arrrow is rendered as ":-" in ASCII
friendly manner, conjuctions are represented by the comma, ",", and variables
are represented by upper-case letters, yielding the standard Prolog -and
Datalog- notation:
grandfather(X,Y):- father(X,Z), parent(Z,Y).
The above clause is a Horn clause. A clause is Horn when it has at most one
positive literal. A Horn clause is definite when it has exactly one positive
literal. Horn clauses with 0 positive literals are called "goals", Horn clauses
with exactly one positive and 0 negative literals are often called "unit
clauses" and Horn clauses with one positive and any number of negative literals
are usually called "definite clauses" (confusingly). A definite clause is datalog if
it has no functions of arity more than 0 (constants are functions with arity 0) as
arguments to a literal. For example, in the following, [1] is Datalog, [2] is not
(but is Prolog):
s(0). % [1]
s(N):- s(s(N)). % [2]
Where s(N) is a function (possible to determine syntactically because it's an
argument to a lieral). In Prolog parlance, definite clauses are also called
"rules", unit clauses are also called "facts" and goal clauses are also called
"queries".
Now, s(0) is a Prolog and Datalog fact and is just as fine a SQL table, called
"s" and with a single row with one value, "0". Here's a fuller example:
father(bob,john).
father(john,alex).
grandfather(X,Y):- father(X,Z), parent(Z,Y).
That's a Prolog and Datalog program with two "facts" and a "rule". The following
are two queries and their results:
?- father(X,Y).
X = bob, Y = john ;
X = john, Y = alex.
?- grandfather(X,Y).
X = bob, Y = alex ;
false.
Each query starts with "?-" at the command-line and ends with a "." as all
Prolog clauses. Below the query are its results: the instantiations of the
variables X and Y in the query that make the query true. The ";" means there may
be further results. And "false" means there are no more results.
Now, I leave it as an exercise to the reader (you) to figure out how the above
works out with SQL. Keep in mind that father/2 has a clean translation to a SQL
table named "father" with two columns, for example named "father" and "child".
The "rule" for grandfather/2 is probably best represented as a join.
In any case, as you can probably see, we have here a very different language
than SQL, but with semantics that can be seen as, in a sense, being equivalent
to the semantics of SQL. Except, where SQL makes a distinction between "data"
and "queries over data", Datalog only has facts, rules and queries, that are all
Horn clauses and that are all part of the "program database".
So it's not a complicated machinery on top of SQL at all. The only thing I'm
concerned is of the naturaleness of SQL queries generated by the "compiler"
(some kind of transducer, probably). On the other hand, I reckon SQL is only
meant to work as a kind of "relational assembly" and will not have to be seen by
any human eyes except in rare cases. Or that's hopefully the plan.
Edit: note there are many, Many, MANY variants of Datalog with confusingly subtly different semantics. See the book I recommended in my comment to juki, below. Personally, I get lost in the variations pretty quickly...