Show HN: A Web-to-RSS Parser in Common Lisp
bitbucket.org
bitbucket.org
(load "~/quicklisp/setup.lisp")
(ql:quickload '(:datafly ;; for database access
;; WEB SERVER:
:hunchentoot ;; for providing a web server
:cl-who ;; for building the HTML output
:parenscript ;; for the avoiding of the horrible JS syntax
:smackjack ;; for AJAX requests
:lass ;; for building the (S)CSS styles
You should not use Quicklisp to silently install dependencies as part of your normal
operation sequence, because:+ It imposes its own view on how libraries are retrieved, stored, managed and updated.
+ It is vulnerable to man-in-the-middle attacks.
+ Used in this fashion, it effectively becomes part of your program. The more people that do this, the more ingrained it becomes.
+ Plenty of us do not use Quicklisp at all, preferring other management schemes.
A better scheme for your project would be to have an install-deps-via-quicklisp.lisp script file, that pulls down and installs the dependencies via Quicklisp, for those that want to go that route and also a README that lists all the dependencies (associated github repos or homepages) . That way you decouple normal execution from dependency installation and those of us that do not use Quicklisp are satisfied.
+ It parses source.txt files from quicklisp-projects and builds its own database. Then it can go out and fetch those projects directly [no curation] from their homepages over https/git/hg and so on.
So, quicklisp-projects (https://github.com/quicklisp/quicklisp-projects)
I find to be pretty useful as a centralized index of CL
software, and my thing piggybacks on top of.
+ Alternatively, you give it a Github/Bitbucket/Gitlab/... URL
and it will fetch and register the project for you.Finally, for every project that it has installed/registered, either via following the source URLs from quicklisp-projects or directly, it can automatically check if upstream has been updated and fetch those updates with optional merge/rebase.
This matches the way I work, which is mainly on top of repos, and it also gives me direct access (no curation) to the projects without having to worry about Xach's curated repo being compromised or quicklisp client not doing HTTPS/certificate verification / checksum checks.
Do you see many Python projects being distributed as virtualenv tarballs which quicklisp bundles may be seen as the equivalent to?
Let me present a sane scheme for dependency management, in terms of software distribution:
+ Create an ASDF system definition.
+ Create a README and list all dependencies with links to the code, version requirements (if any) and other notes/gotchas. This is useful to have regardless.
* Optionally, create a install-deps-via-quicklisp.lisp script that pulls down given dependencies via quicklisp.
* Optionally, create a quicklisp bundle.
Quicklisp (and the way quicklisp does things) should never be forced on the user. Don't get me wrong, it simplifies dependency management and has made things easier for newcomers but it compromises on other fronts and is not (and should not become) the standard way of managing dependencies in CL-land since a lot of its design choices do not mesh well with many real-world scenarios. Always leave a fallback.
(loop for system in (list "sxql" "cl-ppcre" "dexador" "clss" "plump" "plump-sexp" "datafly" "xml-emitter" "local-time" "hunchentoot" "cl-who" "lass" "smackjack" "parenscript") do (asdf:load-system system))
I get a load of warnings about "redefining functions". Rather annoying...Here is an example (example.asd)
(defsystem :example
:name "Example"
:description "Example."
:author "foo@bar"
:serial t
:license "BSD"
:version "1.0"
:depends-on (:sxql :cl-ppcre :dexador :clss :plump)
:components ((:file "packages")
(:file "file1")
(:file "file2")
(:file "file3")))
You don't need to specify .lisp extension for files.
Your packages.lisp may be: (in-package #:cl-user)
(defpackage #:example
(:use #:cl #:sxql #:cl-ppcre)
(:export #:*global-symbol*
#:exported-function))
Then you can load your system through ASDF, including all its dependencies,
like so: (asdf:oos 'asdf:load-op :example)
This assumes that example.asd is inside a directory that can be found in: asdf:*central-registry*I added the list of packages to the README and I'll think about a good way to use ASDF without QL and/or usability loss later. Thank you though!
Quicklisp does not replace ASDF (it uses ASDF internally).
Seriously, anyone who's not taken a look: take a look.
Is Common Lisp still worth considering?
I think so. The more I use it, the more I realise how well-put-together it really is. There are numerous places that annoyed me once, before I understood them, and which I now appreciate (pathnames, logical pathnames and Lisp-2 all jump immediately to mind, but there are a few others). The only thing I don't like is the upcasing, but by setting PRINT-CASE appropriately I rarely have to notice that.
The ecosystem is actually quite wonderful these days. Quicklisp is a godsend — xach deserves an award.
All in all, Lisp feels like a well-engineered solution. Not perfect, but really damn good — and better than anything else I've used.
Leaving this aside (after all, it might or might not be a matter of taste), it totally is. The Common Lisp ecosystem is pretty well alive, despite of the rise of Racket and Clojure. You should really give it a try one day. :-)
Clojure is nice. I like the emphasis on functional programming, and a very clean set of basic data structures and operations on them. Plus lots of great libraries.
Nonetheless, being on the JVM seems both a blessing and a curse. And I would love if it was a lot more performant. Clasp is tempting: https://drmeister.wordpress.com/2015/11/23/why-common-lisp-f.... I wish concurrency was better.
The funny thing is that chrome for cellphone already has something like that with the "make this website mobile friebdly".
Anyway, nice project!
THE FASCINATOR: that cheeky hacker Aaron beat you to this ~ http://www.aaronsw.com/2002/html2text/html2text.py and http://www.aaronsw.com/2002/html2text/ ... old, interesting to see if it's still usable after almost 15yrs.
Latest code at: https://github.com/aaronsw/html2text
Opera used to have really great settings (something like 20 different presets) for overriding CSS. Really made things a lot more legible and was a quick fix for broken websites. Sometimes I still use Lynx.
Or a way to programatically parse websites, process with a bit of code, to view sites off-line. Really interesting piece of code.
When trying to create feeds from webpages in the real-world, there are plenty of pitfalls involved, for example: handling JavaScript content, graceful retries, bypass IP blocks, throttle & rate-limit requests, accessing public social media (Facebook, Twitter, Google) etc.
At Feedity (https://feedity.com), we've developed our own little system over the years using .NET (C#) and node.js, with a bunch of tweaks and optimizations, for generating custom feeds from public webpages.
EDIT: And while I'm at it, are his/er packages what one usually uses in common lisp for web development?
I simply scan through the headlines. Typically business news is repetitive enough that you can get a good feel for how "hot" a topic is, and then I just read one or two pieces on it (usually from the "best" writers) and discard the other 50.
Even headline-only feeds are useful in that regard - they're just signal boosters.
On the other hand, I tend to read a fair number of personal/tech blog posts in their entirety - around 20 a day, over breakfast.
Plus keeping up with podcasts which tend to be published less often than articles.
One thing on this approach: you are screwed if the website isn't using well structured HTML. Having to write rules for individual websites will become cumbersome and will break if they decide to change their markup.
I keep track of a few websites off the command line by doing a checksum of the HTML page and then doing a diff. It doesn't work if the HTML includes a timestamp but it does work for the websites that I am watching.
Quicklisp is par for the course with Common Lisp nowadays, and the dependencies all look fairly normal to me.
If you were going for production code, I'd throw Huchentoot behind Nginx.
No - but there are no packages one "usually uses in common lisp for web development". There are choices for every aspect and one is expected to make an informed decision or simply write yet another implementation of the same functionality if preferable in the specific project.
Of the packages mentioned I only use cl-who regularly; I have a modified hunchentoot for dev work and my own css generator. Using a package for handling ajax requests never occurred to me because for all my projects a 5 line macro did the job.
I'm curious, if folks don't use an RSS reader, how do they find articles and follow sites they're interested in?
Certainly the stop-reading-those-websites option is easier to setup than an RSS reader. Maybe that's the "solution" most people adopt?
Although that added efficiency just means I trawl HN for longer...
+1
https://www.spinn3r.com/blog/2016/11/09/The-Death-of-RSS-Lon...
And this kind of codifies my point even though a number of people in the comments here were calling me crazy ;)
It seems the KiTTY website does have an RSS feed at http://www.9bis.net/kitty/data/rss/rssen.xml