Prolog Web Applications (2016)
metalevel.at
metalevel.at
Prolog is ideally suited for web applications: HTML and XML documents are readily represented as Prolog terms and can be reasoned about very conveniently and efficiently.
For instance, let's use Scryer Prolog to fetch all article titles from the HN front page.
Put the following in hn.pl:
:- use_module(library(http/http_open)).
:- use_module(library(sgml)).
:- use_module(library(xpath)).
:- use_module(library(format)).
Then consult the file with: $ scryer-prolog hn.pl
and then post: ?- http_open("https://news.ycombinator.com", S, []),
load_html(stream(S), DOM, []),
xpath(DOM, //a(@class="storylink",text), E),
portray_clause(E),
false.
Yielding: "Apple announces it will switch to its own processors for future Macs".
"macOS Big Sur Preview".
"Most employees of NYT won\x2019\t be required back in physical offices until 2021".
etc.
Scryer Prolog is a modern Rust-based Prolog system that conforms to the Prolog ISO standard and provides several libraries for web applications: https://github.com/mthom/scryer-prolog. A key attraction of Scryer Prolog is its compact internal representation of strings as lists of characters, making the system ideally suited for the use case Prolog was designed for: Convenient and efficient text analysis, which is also required in many web applications.Here is a short video about this topic:
https://www.metalevel.at/prolog/videos/web_scraping
Enjoy!
How would I run this across multiple cores? Handle exponential backoff for retries? Measure the response time for each submission? Store the result in a database?
For instance, library(format), library(http/http_open) and library(xpath) that are used in the example I posted are all written in Prolog, using at most a few very basic and general building blocks that are implemented in Rust:
https://github.com/mthom/scryer-prolog/blob/master/src/lib/f...
https://github.com/mthom/scryer-prolog/blob/master/src/lib/h...
https://github.com/mthom/scryer-prolog/blob/master/src/lib/x...
To run the example across multiple cores, we need more support from the underlying Prolog engine, for instance to run multiple threads. In my view, this is a key aspect of designing a Prolog system: deciding what must be provided by the engine, and what is built on top of it. Scryer Prolog is still in its early stages of development. I expect it to provide support for multiple threads and interfaces to external databases in a few years at the earliest. This also depends on how many contributors the project will be able to attract.
As for the database, this is where prolog shines: prolog is the database. You store data in the internal facts database.
Concurrency is another prolog strength. For example: https://www.swi-prolog.org/pack/list?p=spawn
See also this very nice talk: https://www.google.com/url?sa=t&source=web&rct=j&url=https:/...
As for retries, I have only a blurry idea of what that's about, but from my limited understanding I suspect it's perfectly doable.
Would the internal facts database as storage be as efficient as a typical database, say SQLite or Postgres? What I mean by "efficient" is: Would it consume similar amounts of disk space or memory and would it be as fast in responding to queries with results?
Prolog huge data munging is a thing though. Check the "semantic web" and related for examples.
Prolog can represent HTML elements idiomatically as nested terms. Unlike in Ruby, Lisp, and Lua you can pattern match and manipulate these terms using unification. I'm surprised you don't acknowledge this as at least as powerful. In fact it's more powerful.
I understand if this is too hypothetical for you. The other point still stands, that Prolog's native data model allows you to represent tree-shaped (and DAG-shaped) documents just as naturally as Ruby, Lua, etc.