SBCL: A Sanely-Bootstrappable Common Lisp (2008) [pdf]
research.gold.ac.uk
research.gold.ac.uk
SBCL (16 days ago) https://news.ycombinator.com/item?id=47140657 (107 comments)
Porting SBCL to the Nintendo Switch https://news.ycombinator.com/item?id=41530783 (81 comments)
An exploration of SBCL internals https://news.ycombinator.com/item?id=40115083 (106 comments)
Arena Allocation in SBCL https://news.ycombinator.com/item?id=38052564 (32 comments)
SBCL (2023) https://news.ycombinator.com/item?id=36544573 (167 comments)
Parallel garbage collection for SBCL [pdf] https://news.ycombinator.com/item?id=37296153 (45 comments)
SBCL 2.3.5 released https://news.ycombinator.com/item?id=36107154 (31 comments)
Using SBCL Common Lisp as a Dynamic Library (2022) https://news.ycombinator.com/item?id=31054796 (67 comments)
etc
(this was mentioned below but repeated here)
Am I reading this right that people can (and do??) use images as a complete replacement for source code files?
Smalltalk does have standard text source file format, but that format is best described as human-readable, not human-writable. The format is essentially a sequence of text blocks that represent operations done to the image in order to modify it to a particular state interspersed with "data" (mostly method source code, but the format can store arbitrary stuff as the data blocks).
One exception to this is GNU Smalltalk which is meant to be used with source files and to that end uses its own more sane source file syntax.
"When you use a browser to access a method, the system has to retrieve the source code for that method. Initially all the source code is found in the file we refer to as the sources file. … As you are evaluating expressions or making changes to class descriptions, your actions are logged onto an external file that we refer to as the changes file. If you change a method, the new source code is stored on the changes file, not back into the sources file. Thus the sources file is treated as shared and immutable; a private changes file must exist for each user."
1984 "Smalltalk-80 The Interactive Programming Environment" page 458
~
The image is a cache. For a reproducible process, version and archive source-code.1984 "Smalltalk-80 The Interactive Programming Environment" page 500
"At the outset of a project involving two or more programmers: Do assign a member of the team to be the version manager. … The responsibilities of the version manager consist of collecting and cataloging code files submitted by all members of the team, periodically building a new system image incorporating all submitted code files, and releasing the image for use by the team. The version manager stores the current release and all code files for that release in a central place, allowing team members read access, and disallowing write access for anyone except the version manager."
Images are not replacements of source code files. Images are used in addition to source code files. Source code is checked in. Images are created and shipped. The image lets you debug things live if you've got to. You can introspect, live debug, live patch and do all the shenanigans. But if you're making fixes, you'd make the changes in source code, check it in, build a new image and ship that.
Perhaps further back: before change sets, before fileOut, before sources and change log ? There's a lot of history.
I wonder if the Digitalk Smalltalk implementation "has objects inside that go back to 1977".
but, doesn't building a new system image involve taking an old/existing image, adding/merging all the changes, and then release new image and sources file from that?
in other words, the image is not recreated from scratch every time and it is more than just a cache.
what is described there is the process of source management in the absence of a proper revision control system. obviously when multiple people work on the same project, somewhere the changes need to be tracked and merged.
but that doesn't change the fact that the changes first happen in an image, and that you could save that image and write out a new sources file.
> image is not recreated from scratch every time and it is more than just a cache
Yes, some vm & image & sources & changes can be taken as the base implementation for development purposes -- a persistent cache.
The state of whatever IDE tools were in use will be saved -- is that what makes you say "more than just a cache"? If I sleep a windows desktop is that more than just a cache?
> changes first happen in an image
What if I write a plain-text source code file using Notepad, and use Smalltalk file handling and byte code compilation and command-line argument handling (packaged in the image) to write the result of a computation to stdout (and quit the image without saving)?
yes, so basically what i meant here is that a cache just stores data, but it doesn't store the whole application.
this is significant in that i can shut down an application (say my webbrowser), then i can upgrade it to a new version, restart and and the application will reinitialize itself and load data from the cache, but now i have a new version of the application.
whereas if i put my laptop to sleep, or better yet, hibernate, then the whole state of the laptop is frozen in place, and i can't do anything to it until i run it again. same is true for smalltalk images.
What if I write a plain-text source code file using Notepad, and use Smalltalk file handling and byte code compilation and command-line argument handling (packaged in the image) to write the result of a computation to stdout (and quit the image without saving)?
you could be doing that, but then you would use the image as your IDE and runtime environment, but not building the actual application in your image. so you wouldn't using what i have been taught is the traditional way of doing smalltalk development.
i am not trying to be pedantic here. it does not matter either way. i just find the smalltalk image approach interesting because it forces you to think about software development in a different way.
this mattes to me because i am working with a web development platform (written in pike) that uses a similar approach. albeit more by accident than intentional. the developers of the platform added support for programmable objects that are stored in the platforms database. these objects can change the behavior of the platform itself, like plugins, but because they are stored in the database they can be changed at runtime, like a smalltalk image. and all the same implications for doing that apply here too. the database becomes more than a cache. and in theory the whole platform could be rewritten such that almost all of its code is stored in the database and only a small bootstrapping system needs to remain outside. this is simply made possible because pike can load and update code at runtime and code changes can be applied without restarting, just like smalltalk.
the downside of the image approach is that it makes upgrading the base image harder, because there is no clear distinction between the base image and any user added changes. i kind of have to take extra steps to pick out my changes and apply them to a new image.
it would be interesting if that process could be improved. it probably would require some compartmentalization just like an OS where i have the base OS, my home directory and the system configuration. i can take a disk image, upgrade the OS and the rest still works. it would be nice if upgrading pharo for example would work the same way.
btw: thanks for the email. i have to ask, how did you manage to reply to a comment more than a month old. normally the reply function is disabled on comments that are 14 days old.
Is that in-conflict with a reproducible build process or can we have both.
> makes upgrading the base image harder, because there is no clear distinction between the base image and any user added changes.
We've been keeping "user added changes" in external files? (Plus the changes file.)
Port "user added changes" from the source code archive to each vendor release.
(btw: I ask nicely and don't abuse this small kindness.)
in part the question is what makes a build reproducible. what actually needs to be reproduced? the point of a reproducible build is that a version of source code always produces the same binary.
how do you make reproducible builds in smalltalk? reproducible builds depend on what goes into the build process. so they depend on the compiler and build tools. in smalltalk those are all in the image, and the question is then what happens when i load code into an image. am i getting that right? i am not so familiar with the details here, but i would guess that it depends on how smalltalk compiles the code and how the import process deals with timestamps and source paths, etc.
however if i work on my source code within the image and i share the code by making a copy of the image then the image is the source and the binary and there is nothing to reproduce. your copy and my copy of the image are going to be identical until one of us makes changes to the image.
i'd be curious to learn more here. outside the smalltalk world my editing tools do not affect the reproducibility of the builds of my code. in smalltalk. just getting a new version of the code browser would change the build, wouldn't it? how do you track that or keep that separate?
Port "user added changes" from the source code archive to each vendor release
right, but that's the "wrong" way around from the perspective of a desktop. i don't need to port my code to new versions of VS Code or vim or which ever tools i use to develop. only smalltalk forces me to do that. so i don't mean distinction in the file structure but distinction in the code dependency.
(i would never have considered to ask for being allowed to make a late reply, especially in this case. it is unlikely that anyone else is going to see our conversation. we could have just continued over email. but hopefully we can dig out some worthwhile details that not only me but anyone searching can learn from)
Outside the smalltalk world could getting a new version of the code browser "library" affect the reproducibility of the builds of your code.
(indeed: vanity)
what affects the reproducibility is the compiler, and that's the issue with smalltalk. you can't upgrade the IDE without upgrading the compiler. say if i use pharo, and switch to a new version of pharo then i get a new version of smalltalk, and i can't compile my own code in the image with the old version any more.
based on that i don't understand how you even enable or test reproducibility in smalltalk. i'd like to learn more about that.
(re: vanity. since a while i have wondered how i would go about creating the longest running discussion thread on hackernews. and how long i would be able to keep it going. i think we are off to a good start here. just remember to peek in here once a while to see if there is a response. and if there isn't feel free to poke me by email. (unless we decide that there is nothing to add))
Do you mean because the IDE and compiler are packaged together, so when you get image' it may contain changes to both?
> … i can't compile my own code in the image with the old version any more.
Do you mean because the image' package has compiler' not the old version?
(Probably an abuse of kindness.)
(not as long as we keep to the topic and respond before the time to reply expires)
When packaged together IDE & compiler can't be upgraded separately (without doing the work to in-effect make separate packages).
> … can't compile my own code in the image with the old version
So we could try to compile own-code' in image' with compiler', and we could try to compile own-code' in image" with compiler", but we want to try to compile own-code' in image" with compiler' ?
Not without doing the work to in-effect make separate packages. (Say we could port compiler' to image" but would that mean trying to compile compiler' with compiler".)
And now we're back to what does "traditional way of doing smalltalk development" mean because supposedly "Team/V could forward and backwards migrate versions of Smalltalk “modules” within a running virtual image."
> the point of a reproducible build is that a version of source code always produces the same binary.
Scope: does "a version of source code" just mean own-code or does it mean sources+changes.
in debian, a build is reproducible if the checksum of the resulting package of my build matches that of your build. how do you do that in a smalltalk image? you take the checksum of what? which objects or rather entities do you compare, and how do you compare them, to verify that a build is reproducible? it can't be the whole image, because that is guaranteed to be different somewhere, and thus the image checksums won't match. so how does that really work?
I have been talking about how to reproduce a particular build, not how to verify.
From the same initial state, perform the same sequence of actions, implies arrive at the same final state. Reproducible.
"Retaining your old change logs gives you a record of all the changes you have made to the system. This will prove invaluable when you receive a new release of Smalltalk/V. … In short, back up the image and change log together, and you shouldn't have any problems."
page 285 Smalltalk/V 286 Tutorial and Programming Handbook 1988
so we actually have to go back to this question:
> traditional way of doing smalltalk development
Is that in-conflict with a reproducible build process or can we have both?
to summarize, i was talking about traditional meaning to make changes in an image and then exporting them if needed. and then i was wondering if local changes and the base image can be kept separate, and by that i meant "within the image" so that, when you look at an image, or when you use it, browse the code, you can easily see what are changes and what is the base image. kind of like in an operating system where i can clearly distinguish system files from user files. so effectively that would suggest some kind of ownership markers, or, alternatively, it needs versioning of objects within the image.
i am not suggesting that smalltalk needs to be changed, but i am trying to work out the limitations that affect the workflow. essentially, because this distinction is missing, there is no way to reproduce a build (in your sense of the meaning) without exporting the code and re-importing it in a new image. actually, object versioning within the image would be a really useful addition to smalltalk. it would allow you to easily go back to old versions. of course it would also cause the image to grow, so there would have to be some efficient way to store that. for one, old versions don't need to be kept in memory so they could be in a secondary storage. or put differently, smalltalk could keep track of the changes in the changes file as it does, but in such a way that i can browse and load in old changes from within the image. i don't know if that is possible now.
the image based server system that i talked about earlier btw does have an ownership system and object based versioning. that reveals some interesting potential.
btw this lack of distinction between base image and changes in the image, especially the lack of ownership information i believe is one reason why linux distributions didn't pick up smalltalk as soon as it became available under a free license, because the license could not actually be verified across all objects. especially older objects that had the source purged to save space.
You seem to be asking about something that is in documentation from 1984.
http://stephane.ducasse.free.fr/FreeBooks/TheInteractiveProg...
> i am not suggesting that smalltalk needs to be changed…
Smalltalk has been changed, several times.
"With ENVY/Manager the image becomes a discardable by-product of development, secondary to the repository that stores and coordinates all work."
https://www.google.com/books/edition/Mastering_ENVY_Develope...
"The way ChangeSets are created and managed in Cuis is different from Squeak."
https://github.com/Cuis-Smalltalk/Cuis-Smalltalk-Dev/blob/ma...
i am also not suggesting that smalltalk can't be changed, just that i didn't want to assume that my suggestion is the right way, and has to be done, or else. more like: how about this? is that feasible?
cuis is interesting, there are other branches that i find interesting too, such as strongtalk. it's a pity that they receive even less attention than smalltalk itself.
i wish i had more opportunity to work with smalltalk, but nobody is hiring ;-)
As-before it doesn't need to be.
SBCL (and maybe others) use a "core image" to bootstrap at startup. It's not unheard of for people to build a custom core image with the packages they use a lot from the REPL. It's become less common as computers have gotten faster, and most people use systems like Quicklisp or Roswell to automatically get updates and load from source. Of course the SBCL core image is generated from the compiler source code when building it, and the dependencies are loaded and compiled from source initially, too, so there's still going to be source code files around.
You could, in theory, start with the compiled SBCL image, exclusively type code into the REPL, save the image and exit, and then restart with the new image and continue adding code via the REPL. I really doubt anybody uses that workflow exclusively, though. At the very least most people will eventually save the code they entered in the REPL into a source file once they've debugge it and got it working.
Several JVM implementations, including the ART cousin, .NET, and apparently node.js is getting one as well.
Even though their appear to be file based, the plugins API makes use of a virtual filesystem that allows for managing the code as if it was the image based concepts from Smalltalk, Lisp and other systems like Cedar, Mesa, Oberon,....
Also something that many don't think about, databases with stored procedures.
At some point maybe the system state will be captured into source code for longer term maintenance, but I can totally see the source code being secondary to the current state of the system during exploration.
After all, that's how I tend to treat SQL databases early on. The schema evolves in the live server, and only later do I dump it into a schema creation script and start using migrations to change it.
Ah, that’s a very helpful analogy/parallel that didn’t occur to me. Thank you!
The proprietary implementations are also quite good.
discussion (2023): https://news.ycombinator.com/item?id=38052564
Let me know if you try it out. I would love some feedback (via github)
Is that only a recency bias ? Because I learned Scheme first ? When I try CL I find my mind resisting some things due to purely elegance reasons.
If someone put a gun to my head and asked me for a deliverable quickly I will go with CL to save my life of course, but for my own personal pleasure I will always prefer Scheme.
In fact, I'd say CL is too minimalist, hence CLDR (a bit like SRFIs) and various libraries which form a sort of unofficial extended core (Alexandria, Bordeaux threads, CL-PPCRE,...)
My comment said nothing about language permanence, though I would say that some measure of evolution can sometimes be for the better. I doubt many people would prefer programming in Java 1.4 over Java 21.
Google Flights is an acquisition of a company using Lisp, ITA Software, they even have a Lisp guide.
https://google.github.io/styleguide/lispguide.xml
In Portugal, Siscog used to be a Lisp shop, no idea nowadays.
Then you have the Clojure based companies, where Datomic and Nubank are two well known ones, even if not a proper Lisp, still belongs to the same linage.
I was aware of the company when I was still living in Lisbon, a few decades ago.
There is a lot more as well, of course, but these two are clear examples of Common Lisp being used in 'the real world'.
it _is_ Lisp. Namely lisp-1, vs what one would consider lisp like common lisp would be lisp-2. Difference mostly being that in lisp-1 everything's in single namespace, whereas lisp-2 has more. So, in scheme you cannot have a function and a variable have the same name. In common lisp you can. Other diffs being (syntactically) passing functions and executing them. There are other things, of course, but not that big of a deal. Scheme is simpler and suitable for teaching / getting into lispen. I'd argue it might also be a rather well-equipped DSL.
https://www.grammarly.com/blog/engineering/running-lisp-in-p...
Some companies: https://github.com/azzamsa/awesome-lisp-companies/ (Routific, Google's ITA Software, SISCOG running resource planning in transportation, trading, big data analysis, cloud-to-cloud services, open-source tools (pgloader, re-written from Python), games (Kandria, on Steam and GOG, runs on the Switch), music composition software and apps…
More success stories: https://www.lispworks.com/success-stories/
I myself run web-apps and scripts for clients. Didn't ditch Django yet but working on that.
And this in not something lisps explored much (is there anything at all apart from Racket/typed dialect?), probably due to their dynamic nature. And this is why I dropped lisps in favour of Rust and Typescript.
Links to Coalton and related libraries and apps (included Lem editor's mode and a web playground): https://github.com/CodyReichert/awesome-cl/#typing
Sorry but I don't compare to C anymore, I want the same safety as in Rust or Typescript: exhaustive checks, control-flow type narrowing, mapped types and so on. Some detection at compile time is not enough, since there is a way to eliminate all type errors I want to eliminate them all, not some.
SBCL does a fine job in detecting type mismatches within the frame of ANSI Common Lisp, not Haskell. While I would agree that a strict type system eases long term maintenance of large systems, for "explorative computing", proof-of-concepts, RAD or similar that tends to get in the way. And if such proof-of-concept looks promising, then there is no shame in rewriting it in a language more suitable for scale and maintenance.
> for "explorative computing", proof-of-concepts, RAD or similar that tends to get in the way
I would even argue that its better to have typed system even for POCs, because things change fast and it very often leads to type errors that need to be discovered. At least when I did that I often would do manual tests after changes just to check if things work, with typing in place this time can also be minimised.
Uh, isn't that exactly what happens with runtime type checking? Otherwise what can you do if you detect a type error at runtime other than crash?
In C the compiler tries to detect all type errors at compile time, and if you do manage to fool the compiler into compiling badly typed code, it won't necessarily crash, it'll be undefined behavior (which includes crashing but can also do worse).
No, it raises an exception, which you can handle. In some cases one can even resume via restarts. This is versus C, where a miscast pointer can cause memory corruption.
I'm curious in what situation catching a typing exception would be useful though. The practice of catching exceptions due to bugs seems silly to me. What's the point of restarting the app if it's buggy?
Likewise, trying to catch exceptions due to for example dividing by zero is a strange practice. Instead check your inputs and throw an "invalid input" exception, because exceptions are really only sensible for invalid user input, or external state being wrong (unreadable input file, network failures, etc.).
Exceptions from type checking are useful because they tell you exactly where something has screwed up, making fixing the bug easier. It also means problems are reduced from RCEs to just denial of service. And I find (in my testing) that it enables such things as rapid automated reduction of inputs that stimulate such bugs. For example, the SBCL compiler is such that it should never throw an exception even on invalid code, so when it does so one can automatically prune down a lambda expession passed to the COMPILE function to find a minimal compiler bug causing input. This also greatly simplifies debugging.
A general reason I look down on static type checking is that it's inadequate. It finds only a subset, and arguably a not very interesting subset, of bugs in programs. The larger set of possible bugs still has to be tested for, and for a sufficient testing procedure for that larger set you'll stimulate the typing bugs as well.
So, yeah, if you're in an environment were you can't test adequately, static typing can act as a bit of a crutch. But your program will still suck, even if it compiles.
The best argument for static typing IMO is that it acts as a kind of documentation.
> It finds only a subset, and arguably a not very interesting subset, of bugs in programs. The larger set of possible bugs still has to be tested for, and for a sufficient testing procedure for that larger set you'll stimulate the typing bugs as well.
I agree that having your program be properly typed is a minumum requirement for it to even be possibly correct. So why would you not check that requirement statically? Skipping it is akin to not checking the syntax at compile time, instead preferring to crash at runtime due to invalid syntax. Or worse, throwing an exception. And then you're supposed to catch the "missing parenthesis exception"?
If you're building a bridge, presumably you'll calculate the forces and everything, before putting it together and verifying it actually holds for those forces. Likewise, if I specify that my function f returns a positive integer, why would I not let the compiler check that assumption statically? Are you really saying you'll write a test like (assert (> (f) 0)) or whatever? Seems like a huge waste of time building the app and starting up the test suite, when you could just as well just have stopped the build when the compiler saw you were returning the wrong type.
> For example, the SBCL compiler is such that it should never throw an exception even on invalid code
What does this even mean? What does it do if you run (/ 1 x) where x = 0?
> So, yeah, if you're in an environment were you can't test adequately, static typing can act as a bit of a crutch. But your program will still suck, even if it compiles.
This is a ridiculous argument, and is much closer to "just don't do the bad things" than what I said. Also note that I never said "don't write tests". You're making a complete strawman by pitting type checking against testing. It's two complimentary methods of verifying correctness, and the type system is a quicker and more accurate way of catching bugs, so moving all logic that you possibly can to the type system and away from tests is a huge win in efficiency and correctness. That doesn't mean skip testing, just like checking the syntax of your program statically doesn't mean skip testing.
Just because I'm wearing a seat belt (type checking) doesn't mean I'll drive recklessly (skip testing). Likewise, just being a careful driver (writing tests) doesn't mean you should stop wearing your seat belt (type checking).
"Structure and Interpretation of Computer Programs"
https://web.mit.edu/6.001/6.037/sicp.pdf
See https://planet.racket-lang.org/package-source/neil/sicp.plt/... as well.
If you're interested in LeetCode, Racket is one of their accepted languages.