Prolog for data science
emiruz.com
emiruz.com
Datalog is arguably the minimal core logic programming, similar to what the lambda calculus achieves for functional programming. Unfortunately, it has been forgotten outside of database and query processing realm. A resurgence has happened in recent years, as PL researchers and also industry have discovered the virtues of datalog (e.g. Flix, DataFun). My own attempt at making this more widely known is here https://github.com/google/mangle, a language from the datalog family and its implementation as a go library.
As the example shows: plain "rules" (or: plain datalog) is rarely enough to capture everything that one wants to express: the question then is, how to combine a pure declarative "kernel" with more general purpose programming (e.g. mapping a list).
PROLOG offered one answer, already in the 1980s, but I fully reject it: the fact that the writing a program in the wrong order with negation and recursion makes it non-terminating is not something we'd want everyone to deal with. Datalog with stratified recursion is somewhat better, as "layers of rules" is a concept that is easy to understand.
In mainstream programming languages, the possibility of writing non-terminating programs also exists, but is rarely an issue. That is why I believe a good combination of declarative and general-purpose has to make it really easy to recognize which parts of a program are in the declarative, terminating, safe kernel and which parts require more attention from the programmer.
The kind of standardization that happened for SQL and PROLOG certainly helped spread its use, but it is a very differently world today. Developers can easily do their own take on DBs' data warehouse by serving from memory or existing DBs or files.
If you do not see it as a programming language, but a way to think about computation, then we are in the ballpark of "rule engines:" there are of course innumerable implementations of things that are called rule engines. Like the post, "rules" make knowledge explicit, but we wouldn't even think about the possibility that all the folks who wrote or use these rule engine implementations use datalog syntax. It is more the semantics, structuring the problem as facts and rules, that counts.
Of course having a common syntax helps and matters in getting the message out: that there is a good foundation. But how to add aggregation or user-defined functions is not settled (it is also not settled for SQL, many vendor-specific extensions, syntaxes...) and I think today's business world does not provide much incentives for agreeing and standardizing.
Maybe academia will be able to help over time, by teaching newer generations who will then pick standard syntax for their next PL because they are familiar. For academia to be interested in an applied PL topic, it has to be reasonably formal, derivable from first principles and teachable.
Datalog is very basic and everyone needs aggregations and structured data so Mangle also supports aggregations, structured data and also some function calls. When you use these it is clearly no longer datalog, but it is easy to see what part is datalog and what part is "more."
To qualify this a little: Mangle does not provide much for mutability (neither the "spec" which is a bit implicit, nor the implementation) so if you want to make a real DB with inserts or an RPC interface you have to code that yourself.
I find having a readable source file and running queries is good for playing around and also may cover many use cases for small DBs with static or slow changing data, or configuration.
[1]: https://quantumprolog.sgml.io/bioinformatics-demo/part1.html
Anyone interested could also take a look at Popper (https://github.com/logic-and-learning-lab/Popper) or this overview of the first 30 years of ILP (https://arxiv.org/abs/2008.07912)
I made a few contributions to Aleph as part of my PhD on doing transfer learning in ILP and really enjoyed working in Prolog.
Jokes apart, what makes me able to detect that kind of applications of LLM is that I have been looking how to combine LLM with rule based systems, using statistical methods, so I analyze anything that smells like that.
Gelman et al have written a lot about this, and they have a proposed general workflow [1]
[1] Bayesian Workflow https://arxiv.org/abs/2011.01808
There is a certain clarity of purpose declarative code has that I find really pleasing.
Also I'm lazy and I prefer if the computer thinks for me.
If so, would sorting the segments by start position first make it much easier? Start at the far left, go right merging the overlapping ones until there's one which doesn't overlap?
(Please don't name the graph axes X,Y and then use (X,Y) in the code for things which aren't X,Y data points - anything wrong with sticking to (I1, I2)? Why use "R" for span length? And what's "A"? And nitpick "(I1, I2)" is not a tuple in Prolog, it's a term and "I1-I2" is a more idiomatic and fewer parens term which can be used in the same way. It's what the imported pairs library uses, for example - https://www.swi-prolog.org/pldoc/doc/_SWI_/library/pairs.pl )
In light of all that, making decisions by gut feelings or intuition is, if nothing else, at least honest, and probably just as good an approach as anything else.
If we were all to iteratively acknowledge our interests to a greater degree then we could probably learn a lot more from data but I think there is a strong message in our society that self interest is wrong and so people dedicate great energy to pretending that their self interest does not exist.
I went down on this thought once, and I became antiscience. Reasoning-wise I only trust what I can understand, simple highschool level reasoning. Anything that I can't understand, and is more sophisticated than a highschool level reasoning I don't trust. I mean, sure, there are facts that need complicated statistical methods. But I can't check them, I don't trust authority, (and probably they are more often wrong than right anyway, because they need results or they starve to death), so I reject them/I am ambivalent.
Like what? There are ideas that need causal reasoning, or trust that it exists behind things you don't fully understand. Statistics, especially anything higher order, end up being basically just rhetoric, outside maybe of some very narrow claims.
If someone tells you you need a bunch or statistics to understand something, they're almost certainly trying to persuade you without having a strong enough logical footing to just explain themselves.
And science isn't a religion or a position, you can't really be for or against it. That should be a starting point.
Causal mediation analysis is similar if you want to understand what assumptions you really need to make to talk about the mechanisms through which a treatment effect acts on the outcome.
If you're talking about optimizing around some agreed upon thing, sure I can understand the idea of using some more complex statistical analysis (though I might question the incremental value).
To understand when a simple answer is unlikely to be correct, you do benefit from understanding the mathematics. There’s a reason that people with doctorates in statistics spend 8 to 10 years in school to learn how to contribute to it.
https://github.com/patricktrainer/csuite-saas-metric-generat...
It’s public perception of modern science that is that way (in large part because of politicians).
Still bad, but a different kind of bad.
Why do you want to do a linear regression on random segments?
Author talks about reasoning, but I think it would make the point clearer if there was a section that reasoned about the result on FTE.
Why are there a white upward segment and a white downward segment? Inside the blue segment there is a clear downward segment, about the size of the red downward segment. Why did it not get its own color?
It may be worth reading the article fully -- why random segments and the logic of it is explained fairly well I think.
>Author talks about reasoning, but I think it would make the point clearer if there was a section that reasoned about the result on FTE.
Same as above. Have a read of the comments around the implementation.
>Why are there a white upward segment and a white downward segment?
There are no white segments -- those are gaps.
> Inside the blue segment there is a clear downward segment, about the size of the red downward segment.
I guess you mean the last graph? I think if the isotonic regression was plotted on the graphs, it'd be clearer why that segmentation makes sense. But in brief it is about how well an isotonic regression fits in the segments as opposed to how it may visually seem -- the two can come apart.
I already know:
* Strand
* http://www.call-with-current-continuation.org/fleng/fleng.ht...
* Racket's implementation of datalog
TIA.