Scallop – A Language for Neurosymbolic Programming
scallop-lang.org
scallop-lang.org
I really love the concept. This isn't just differentiable neurosymbolic declarative probabilistic programming; Scallop has the flexibility of letting you use various (18 included) or custom provenance semirings to e.g. track "proofs" why a relational fact holds, not just assign it a probability. Sounds cool but I'm still trying to figure out the practicality.
Also worth pointing out that it seems that a lot of serious engineering work has been done on Scallop. It has an interpreter and a JIT compiler down to Rust compiled and dynamically loaded as a Python module.
Because a Scallop program (can be) differentiable it can be used anywhere in an end-to-end learning system, it doesn't have to take input data from a NN and produce your final outputs, as in all the examples they give (as far as I can see). For example you probably could create a hybrid transformer which runs some Scallop code in an internal layer, reading/writing to the residual stream. A simpler/more realistic example is to compute features fed into a NN e.g. an agent's policy function.
The limitation of Scallop is that the programs themselves are human-coded, not learnt, although they can implement interpreters/evaluators (e.g. the example of evaluating expressions).
In the other hand, these problems are routinely analyzed and solved by differentiable algorithms running on neural net substrates (e.g. you).
https://www.cis.upenn.edu/~mhnaik/papers/neurips21.pdf
https://dl.acm.org/doi/10.1145/3591280
There is a 135 page book on Scallop https://www.cis.upenn.edu/~mhnaik/papers/fntpl24.pdf
Fodor and Pylyshyn's Critique of Connectionism and the Brain as Basis of the Mind https://arxiv.org/abs/2307.14736
> The work uses graphs developed using methods inspired by category theory as a central mechanism to teach the model to understand symbolic relationships in science.
https://news.mit.edu/2024/graph-based-ai-model-maps-future-i...
Category theory can be leveraged to make faster theorem provers (making complex symbolic reasoning practical at larger scales).
Don't ask me how, hopefully someone who studies it will chime in and correct me / expand.
you seem to be more in the know than me :) Please could you just sketch out a few bullets and explain the relationship between Scallop and Lobster and what you think is going on?
Just look at the examples on their website. All 3 are lame and far easier without their language.
It's like publishing that you have a new high performance systems language and never including any benchmark. They would be rejected for that. Things just haven't caught up in the ML+PL world.
It's not about performance, but safety.
Making safe decisions becomes exponentially more important as ML / agents evolve, to avoid "performant" but ultimately inefficient/dangerous/wasteful inferences.
I may re-evaluate now, thinking of smoother LLM integration as well as differentiability.
Has anyone here used Scallop for a large application? I ask because in the 1980s I wrote a medium large application in Prolog and it was a nice developer experience.
Very pleasant branding though. Great work! :)
So it's intended to combine nn reasoning and logical reasoning cleanly.
Oh, boy, it's written in Rust!
A neural network (PyTorch) detects objects and actions in the image, recognizing "Jim" and "eating a burger" with a confidence score.
A symbolic reasoning system (Scallop) takes this detection along with past data (e.g., "Jim ate burgers 5 times last month") and applies logical rules like:
likes(X, Food) :- frequently_eats(X, Food).
frequently_eats(Jim, burgers) if Jim ate burgers > 3 times recently.
The system combines the image-based probability with past symbolic facts to infer: "Jim likely likes burgers" (e.g., 85% confidence).This allows for both visual perception and logical inference in decision-making.
So when would you use symbolic programming? To generate quality data for the neural network. For example, maybe the neural net reports it read the speed limit to be 1000 km/h on a sign because of someone's shenanigans. A symbolic programming aid which knows potential legal limits will flag this data as potentially corrupt and pass it back to the network as such allowing the neural network to take more sensible decisions.
Why is this a language and not just some say, Java/Rust library?
It's interesting but doesnt seem like fundamentally anything new.
By the way I wish there were more real-life examples, both basic and advanced to show what it may be especially useful for, maybe even compare it to other languages like Prolog. I expected the tutorial to have examples for what "neurosymbolic" means, because I am not entirely sure what it means in practice.
It seems like schizo ramblings to me. But I'm sure there's some merit to it.
But - it's time to do the tutorials and try and see.
https://en.wikipedia.org/wiki/Fibonacci_sequence
This one was easy to spot and would have been easy to get right. Makes me wonder…
> Many writers begin the sequence with 0 and 1, although some authors start it from 1 and 1[1][2] and some (as did Fibonacci) from 1 and 2.
> Many writers begin the sequence with 0 and 1, although some authors start it from 1 and 1[1][2] and some (as did Fibonacci) from 1 and 2.