Blog with Markdown and Git, and degrade gracefully through time
brandur.org
brandur.org
After a year or so I migrated to a 5 USD droplet on Digital Ocean (back then GH Pages didn't offer https for custom domains) and integrated with Github webhooks to automate the deployment when pushing new markdown to the main branch.
Over the time it indeed started to degrade. The build time takes almost a minute, but after building the website is just a bunch of static html pages.
Nowadays it is annoying to write new posts because I like to write locally and refresh to browser to check if it's looking good or not. So I would say it degraded for me but for the reader it's still as fast as it was when there was just a couple of posts.
I thought about migrating the blog to something else, but because I used some custom markdown extensions for code highlight and other things, it would be painful to migrate all the blog posts. So I've been postponing it since 2019.
Then I learned about how hard it is to use a cache.
to be honest jekyll still serves the purpose really well (especially for the reader)
the site is still sitting on an inexpensive entry-level cloud server, it's relatively fast, serves 300k+ page views every month and it give me some passive income
i'm pretty sure that if one uses the framework properly it can still hold on very well
i think what left me with a mixed feeling about static site generators is that it is constantly holding you back whenever you think about growing your website. i ended up building a django api that runs on the same vpc as the blog itself, and i used it to expand some features, like reading from google analytics the total page views for a given blog post, or consuming the disqus api to list the latest comments on the home page. this kind of thing.
managing the posts in static files is quite challenging as the number of posts grows. at some point you will want to change some info, or add a certain metadata and you will need to write a script to walk the _posts dir and edit the files (maybe there's a better way of doing that :P)
Find all posts related to "django", select the top 3, pin them as the related posts for all django entries? You'd trade out some uniqueness and slightly increased memory usage to gain speed.
(I just got started blogging and found https://github.com/sunainapai/makesite python static site generator worked well for me. Obviously you're probably a bit far in to switch horses by now.)
Since that time I look for the people and community behind the project, and try to find signs of stability and long-term care. After that I look at open formats rather than open and flexible architecture-chains. For example, I'd rather use my LibreOffice HTML template and simple PHP controller on a more monolithic (but open) platform than connect a bunch of technologies together to create a build process with a bunch of moving, quickly developing, interdependent parts.
Not sure it's the best answer, but it has worked better to use more monolithic software, even blogging software that's been in steady, if slow development since the early 2000s...
I’m currently on Hugo 0.80.0 and who knows, it might take years before I upgrade it.
Could you tell me more about this? I've had a similar idea but the HTML that LibreOffice generates is very gross
Their secret sauce is, effectively, partial evaluation in Docker images. They run the code to detect if any of the changes in a layer have side effects that require that layer to be rebuilt (which invariably causes every layer after to be rebuilt)
I mention this because if I'm editing a single page, I would like to be able to test that edit in log(n) time worst case. I can justify that desire. If I'm editing a cross-cutting concern, I'm altering the very fabric of the site and now nlogn seems unavoidable. Also less problematic because hopefully I've learned what works and what doesn't before the cost of failure gets too large. It would be good if these publishing tools had a cause-and-effect map of my code that can avoid boiling the ocean every time.
I believe the last time I touched make was to fix an exceedingly badly mapped out chain of cause and effect that would consistently compile too much and yet occasionally miss the one thing you actually needed.
Partial evaluation works out the dependencies by evaluating all of the conditional logic and seeing which inputs interact with which outputs.
Make needs that information, but that doesn't mean you need to be the one to catalog it. I found this article pretty useful [1]; in a nutshell, many compilation tools (including gcc, clang, and erlc) are happy to write a file that lists the compile time resources they used, and you can use that file to tell Make what the dependencies are. It's a bit meta, so it made my brain hurt a bit, but it can work pretty well, and you can use it with templated rules to really slim down your Makefile.
(I've tested this with GNU Make; don't know if it's workable with BSD Make)
[1] http://make.mad-scientist.net/papers/advanced-auto-dependenc...
It’s very fast.
I have a mockup blog with ~90 pages and it takes 190 ms to generate the whole site.
But, once you approach two to four or five hundreds, it can quickly add up to a minute or two, making it impractical to say the least.
One solution is to run `jekyll build`, which will just build HTML to a directory, and then just removing old Markdowns and serving those generated HTMLs directly via nginx or something.
I've honestly given up and switched to Ghost, where I don't have to worry about that sort of stuff.
It is really astonishing what power our devices have and how little that power is utilised by many tool chains.
Having the SQLite database made cooking up views really trivial (so I wrote a plugin to generate /archive, another to write /tags, along with /tags/foo, /tags/bar, etc). But the process was very inefficient.
Towards the end of its life it was taking me 30+ seconds to rebuild from an empty starting point. I dropped the database, rewrote the generator in golang, and made it process as many things in parallel as possible. Now I get my blog rebuilt in a second, or less.
I guess these days most blogs have a standard set of pages, a date-based archive, a tag-cloud, and per-tag indexes, along with RSS feeds for them all. I was over-engineering making it possible to use SQL to make random views across all the posts, but it was still a fun learning experience!
You can still have the sources in a git for quick rollbacks.
What I like about lektor over most CMS solutions is that it is more easily adjustable. Basically Jinja2 and python held together with glue
I run a Django development agency, and I've passed around links to your tutorials more times than I can reliably estimate; hundreds of times :)
It's a great resource, thank you for it.
I’d make a simple static blog generator but then grow tired of some shortcomings, which then led to abandonment.
My current blog has an HTML rendered homepage but all the posts are text files. https://blog.webb.page
This is deeply problematic. We should have dozens, if not hundreds of contenders for "the next WordPress" that leverage Git as a foundational aspect of content management. Instead we have a bunch of Contentful clones (no disrespect to Contentful) with REST and/or GraphQL APIs.
It's bananas that if you search for Git-based CMSes you have NetlifyCMS, and…wait what? Is that all??? Forestry gets mentioned a lot because it's Git-based but that's also a proprietary SaaS. I just don't understand it. Is this a VC problem or a real blind spot for CMS entrepreneurs?
The things a content creator might be more interested in aren't really core git strengths: SEO, publishing schedules, editorial back-and-forth, social media, etc.
Hmm, not really core PostgreSQL strengths either. What am I missing?
There are some more. I know of Grav which I intend to use migrating from NetlifyCMS. It has flat-file storage and Git connectivity. I think it can be installed on my shared hosting provider by just unpacking the thing (I don't have Composer / shell access).
If you want a new CMS, give it a shot. But nobody's made a better free version of Jenkins either. It's hard to do and completely unrewarding/unmonetizeable (as FOSS), which is probably why nobody has done it.
However, insisting it leverage a difficult source code version control system is artificially restricting. Start with an MVP of flat files for content and add a plugin system. Somebody will write Git support, but they'll also add other content backends. I'll bet you a million dollars that a custom SQL plugin that does version control will be preferred over Git by 99% of users. But they may wise up and use an S3 plugin instead. Or maybe all three. They'll have choices, and your project will become more useful.
Don't start your project with an artificial restriction and it will be better for it.
The real reason most websites disappear is a much more human one.
The more plugins you use, the faster entropy kicks in without maintenance.
With WP in general though I definitely agree.
For simple sites, especially where we don't do comments, I've started using Simply Static though to just generate pre-rendered pages. All the benefits of static site builder with all the conveniences of a CMS.
How many sites have you tested it on, have you had any issues? I'm gonna play around with it now but what I'm mostly concerned about now is how you manage the site if you point the web server to the static directory.
I made a prototype at one point where you have your actual WP site in a subdirectory, like example.com/wordpress, which is IP restricted (or behind http auth). Then you crawl that site with wget (or curl) to generate static HTML and to finish it off by search/replacing the HTML files (removing the /wordpress portion). Then you serve the static files to general users.
- WP moved to subdirectory (WP home and siteurl changed in DB)
- Static site served normally from (and generated to) webroot
Just needed to add sitemaps and robots.txt to the config to be included. Workflow is super smooth. Login to example.com/wordpress/wp-admin/, view your changes at example.com/wordpress and when you're happy, publish the site through the plugin.
Curious to see what kinds of issues I might be running into later on (this was just a quick test).
I don't base on number of plugins, I base on complexity personally.
I mostly agree with you, but in my experience most (not all) people who are capable of using a static site generator will do so. Those who aren't comfortable with it use WordPress. I'm not slamming WordPress, just making an observation.
So true.
I have had a blog running since 2006, all was php based back then with (html) posts inside a database. The tooling shouldn't be a problem, I switched systems several times (once every few years).
Currently on Hugo with markdown posts. As long as you treat your migration carefully and take some time to migrate, every tool should suffice. It's mostly about human effort and human error when things get lost.
[0] Also, thanksnotreally for cherry-picking and putting words in my mouth with your use of `!M$`, screw you too.
Personally, I have a different take than the OP, although I'm not sure this is really a disagreement. I would argue open data formats are more important in this context. If I write text in Markdown, it doesn't matter whether I'm using Emacs or a closed-source, proprietary commercial editor; the text isn't tied in any meaningful way to the editor. Likewise, a Git repository hosted on GitHub isn't tied in any meaningful way to GitHub. It's... a git repo.
As long as GitHub isn't doing anything that I particularly object to -- and "oh woe, it is owned by Microsoft" is not an objection I share -- there's a pretty good case for continuing to use it. I'm taking a calculated risk that it's both unlikely to go away anytime soon and to markedly change business direction in an unfriendly way, but if it does? As long as my local copy is up to date, I can re-publish it anywhere and just, well, stop using GitHub.
Not that I disagree with this, but that goes almost directly against "Mak[ing] that source public via GitHub or another favorite long-lived, erosion-resistant host. Git’s portable, so copy or move repositories as you go."
Any self-hosted service or solution is going to not be erosion-resistant by virtue of not being for-profit.
Of course the end-result is easy to move out but the build pipeline is more dependent on GH.
Not that it matters but it could be cool too to pipe SSG's output through a managed VPS just to show you don't need Github to host your static files.
And for-profit services are? Often enough it's the VCs which ruin a service by pressing out money before it inevitably dies because the experience has been degraded to an extreme degree.
My co-author and I use Markdown and Git as the author suggests, and one of the best things is that between simple CI/CD pipelines and effortless scaling of a static site, we don't need to do any technical work so there's no friction on my lifestyle. We've been writing for almost a year now, 4+ posts a month, and 99% of that "work" has just been writing.
Writing with a partner helps a ton as well.
It's been great.
There are some blogs that I realize will never "come back" so long as "everyone is on Twitter these days". Because Twitter is still so much more convenient that blogs (even ones in Markdown + Git).
[1] Okay, not actually random, it was the Google Reader shutdown year. Google Reader provided a lot of convenience to RSS, including social media-like network effects, that almost brought blogs mainstream.
I have a popular Spark blog on Wordpress (https://mungingdata.com/) and can relate to your sentiment that any inconvenience can hold up the writing process. Your post is motivating me to streamline my publishing process.
https://gohugo.io/hosting-and-deployment/
If you have no idea, Netlify is good and with a good free tier.
We've gotten it to the point where we just push new pages via git and it all works. Anything more and you'll find yourself procrastinating ever so little at a time...
Typically it sits in the back of a dusty rack for a website hosting vendor (or, rarely, a colo provider) and the gear has long since paid for itself, but also is unmaintainable due to not even having the remotest semblance of a service warranty or service parts (other than what somebody might have ordered as spares years before). If it ever loses power, everything starts back up on boot time and it keeps on chugging, defying the usual laws of computer entropy.
I haven't setup CI pipeline yet.
It true whats said in the post, I used to save links thinking it won't disappear. But most the links, I had saved, it's no more.
Of course I'm not sure what other people are using them for, but in my case GitLab/GitHub covered all my bases pretty well since I wasn't looking for a pelican/hugo CMS and you still host somewhere else (not Netlify).
Netlify's main business is.
Netlify also provides deployment previews for commits and PRs, and a complete CI/CD pipeline that you would have to implement yourself with Github Actions or something else.
The CMS isn't what I need either, but there is so much more than just hosting.
They can also manage your domain/ssl which is nice so I can manage everything related to my site on Netlify.
Lmk if you need any extra details
When you feel you want more, you can link a git repo and point a domain name at it. When you update the repo, it automagically updates the site.
It's my goto for all new projects.
There is great documentation at every step.
Nowadays I'm planning on writing again and have tried a setup that seems to be alright:
- migrate from Pelican to Hugo since the community is more vibrant, responsive and the themes much more mature and well maintained.
- Use a mature, well-maintained documentation-oriented template as base (I based https://jdsalaro.com off Pelican's bootstrap3)
- Create an Obsidian vault out of the posts/ directory to improve note-taking and editing capabilities.
- Host same as before on GitLab
- Deploy via CI/CD copying the posts/ directory to a public repository
- Probably mirroring to GitHub
I'll be able to report back once I've gotten the above fleshed-out :)
I can still easily change the private flag down the road if I were to decommission it (though, unlikely -- I'd just archive it somewhere).
I'm also using Gatsby if anyone is curious (https://github.com/cbeley/beleyblog & https://chrisbeley.com).
Edit: this is easy, in fact: you make a copy of your 'publish' branch on the publishing site, call it 'draft', and then protect it.
I need more than 10 drafts (the current limit in Ponder), but I feel like I'm getting close to a solution that works better for me. I also want to be able to edit across machines (Mac, Chromebook, Phone), so I need to make the app work online.
Anyway...
Most broadly, it's called content-centric networking. Bittorrent is a piece of that puzzle, but too static, with no obvious way to connect disparate hashes together into a single entity. IPFS and Secure Scuttlebutt are groping in the right direction.
There was a project called gittorrent which, as you might guess, was trying to be 'bittorrent for git'. It never really went anywhere, the crew at https://radicle.xyz are looking to revive it and I wish them the best of luck.
What I want is a single handle, such that, if there are still copies of the data I'm looking for, out there on the network, I can retrieve them with that handle. And also, any forks, extensions, and so on, of that root data, with some tools to try and reconcile them, even though that may not always be possible.
That would be really powerful. It would made information more durable and resilient, and has the potential to change the way we interact with it. Like I find typos in documents sometimes, it would be nice to be able to generate a patch, sign it so it has a provenance, and release it out into the world.
When I browse old blogs which still exist, I routinely hit links to other blogs, video, and the like, which are just gone. Sometimes Wayback Machine helps, often it doesn't. This problem can't be fixed completely, when data is gone, it's gone, but we could do a lot more to mitigate it than what we're doing now.
I think for business websites it starts to make sense to use something like Wordpress. The fact that it’s open source is amazing, so you can always self host. But it’s more effort. But you get lots of neat plugins and templates. But it’s more complicated.
Both are great solutions, but my current thinking is ssg for personal, for business maybe Wordpress.
I use a bunch of old Unix tools that will be around forever to make it look nice: http://www.oilshell.org/site.html
The toolchain has changed significantly in 4+ years. It started as literally a shell script invoking Gruber's original markdown.pl. Then I switched to CommonMark, etc.
But the core data hasn't "rotted" at all, which is good. Unix is data-centric, not code-centric.
Previous thread that mentions that Spolsky's blog (one of my favorites) "rotted" after 10+ years, even though it was built on his own CityDesk product (which was built on VB6 and Windows). He switched to WordPress. Not saying this is bad but just interesting. https://news.ycombinator.com/item?id=25675869
Wouldn't switching formats like that break peoples pages? That's one concern I have with using markdown for a site like this. For anything non-trivial, you either end up using HTML (in which case, what's the point) or some dielect-specific feature.
CommonMark is a Useful, High-Quality Project http://www.oilshell.org/blog/2018/02/14.html
At the bottom of this article I note one incompatibility I encountered with the reference CommonMark vs. markdown.pl. But I just fixed those and have been using it for 3 years, and it's been great.
----
The key point is to have a diversity of implementations, so you don't get locked into one, which may go away in 10-20 years. In 10-20 years, I'm very confident that there will be CommonMark renderers available.
(Even apart from the fact that the reference cmark implementation I use is extremely compact C with no dependencies. C will outlive Perl -- the Lindy effect at work! http://www.oilshell.org/blog/2021/01/philosophy-design.html#... )
Not that blog posts are life changing or anything, but having a format that’s durable and reliable seems important.
I lost my blog from college back in 95 when I got my first job and didn’t think to archive my shell account. I wish I had, even as a memento.
One thing I think that factors into it is that age is kind of nice for slowly removing stuff that’s not used from existence. Posts from 20 years ago might not be good to keep around if no one is reading them. No the gradual degradation that comes naturally from new phones, computers, hosts, jobs might be a feature for some people who don’t necessarily want everything around forever.
Normal people need a way to have a permanent place that can't be taken down and doesn't auto-expire with your credit card.
However, that’s pretty lightweight on decisions I had to actually make to publish, and all the alternatives seem to be more involved. I wouldn’t mind migrating off WordPress, but just on the theming side that has a decent chance of involving a non-free theme, making the idea of hosting it on a public repo somewhat of a non starter...
I'd built this for managing notes, but it seems that many many people use it for their websites. (Including me)
Time is a great filter. Yes important stuff gets lost but not each thought or word is worth preserving and not every moment in life needs to be captured and kept for posterity.
I would prefer if more SO answers, more old tweets, more outdated tech blogs, more old unflattering photographs, ... disappear in the void. The death of geocities meant some valuable stuff was lost, but it also meant that much much more crap was lost.
The world does not gain from keeping every scrap, rather the filter of 'did someone care enough to preserve this' adds value as the signal to noise ratio improves over time.
I would still argue that we live in a time where many things are needlessly preserved (not to mention the privacy implications...).
Now to figure out what I can write about.
Though I'm definitely going to use this for internal docs at work probably. Very nice.
Edit: I'm guessing it's this? https://squidfunk.github.io/mkdocs-material/getting-started/
Here's a site using it: https://docs.libretro.com/
[1]: https://blot.im/
[2]: https://paul.af/
Rendered: https://caddyserver.com/docs/caddyfile
Markdown: https://github.com/caddyserver/website/blob/master/src/docs/...
Unofficial How-To thread: https://caddy.community/t/markdown-support-in-v2/6984
Now if you're arguing for putting the website source on GitHub, that's an entirely different matter. GitHub addresses the human factor by being entirely free with no effort required for site upkeep. That's why it's durable, it's not about keeping it as git and markdown.
I think the author's point is to highlight the relationship between technological and human factors. Git+markdown has a better user experience if your user is a developer who uses git every day.
Lately I've been seriously thinking of reimplementing that, as I just can't seem to find a decent static site generator. They all seem to require a bunch of plugins and configuration just to do the bare minimum a blog requires.
I just want to write, add some code or images, have it look somewhat decent and publish it without worrying about getting hacked due to some WordPress security issue.
E.g. the Hugo quick start guide is pretty much just "install Hugo, pick a theme, add your content" and gets you a barebones sequence-of-posts blog: https://gohugo.io/getting-started/quick-start/
I'll give it a whirl tho, thanks!
Maybe you and I have very different ideas of what is required to write a blog, but it sounds like you should be able to use Hugo + the theme of your choice and just go with it. It's /very/ easy, and posts are all written in Markdown.
From what I can see Hugo does seem to have a lot of that these days, so I'll try that.
One immediate issue is the highlighter not supporting my language at work (Delphi/Pascal) but I assume it being Pygments-compatible means I should be able to add that without much fuss.
The theme I use on my website for instance has a shortcode called "fluid_imgs" which handles images pretty well, and I also wrote my own shortcode addition to support embedding images from Flickr more easily. Additionally the theme I'm using integrates MathJax and HighlightJS, which it utilizes in the standard way for Markdown (e.g. in most Markdown implementations you can specify the language after triple backticks to start a code block) by invoking highlighting on code blocks.
Getting what you're asking for should be pretty simple in Hugo with the appropriate theme or minimal effort anyway.
If that's the case, go with that instead of trying to assimilate into workflows around existing static site generators. The world benefits from diversity, even diversity in software. Consider pouring your effort into making OpenLiveWriter work the way you want, if it doesn't already.
[2] https://sqlite.org/locrsf.html
[3] https://sqlite.org/lts.html
[@] a single file, not a bunch of files in a directory
Bonus feature: since it runs nightly it gives me diffs of changes I make to my content, including edits to old posts.
The backups are in this repo: https://github.com/simonw/simonwillisonblog-backup
It's pretty easy to mirror your own git repos, so that's a nice option, and to have stuff hosted on GitHub, your own, and maybe BitBucket as well. But nothing is going to beat having your own backups--which I DO still have, from as early as 2000.
The simplest subset of HTML you can manage coupled with an uncomplicated layout and no external assets is better in terms of longevity.
/info/refs service=:service https://gitlab.com/<user>/<project>.git/info/refs?service=:s... 301
Then if your website is at https://example.com someone can just 'git clone https://example.com'. They can periodically do 'git pull origin master' and you can also push changes that way. There's no problem moving to a different git repo provider (github/gitlab/whatever).
For example, I recently wrote an article about YC's Work at a Startup that contains some interactive visualizations embedded in the markdown file I used to draft the content as Vue components. If I want to cross post, I'll probably need to maintain another version of the markdown that either contains static images of the interactive elements and/or links back to my site (on GitHub pages).
I'm part way through the implementation of showing the commit history for each post and linking it to the commit hosted on sourcehut. It provides another view on the content the same as hosting a gopher or gemini version would: clone it and read the things in the posts folder however you want.
I then scp the file into the pool directory on my server, and re-run a script that calls apt-ftparchive and regenerates the contents of the repo (https://manpages.debian.org/buster/apt-utils/apt-ftparchive....).
My web server hosts that directory with indexing enabled, but I don't use apache for it like most examples do. There's nothing special about it, it's just a directory tree built in a way that apt likes. (https://pkg.kamelasa.dev/). In fact, the entire configuration of the repo is visible there.
There's a step in the middle where I sign the packages with my GPG key, and the public key is available on Ubuntu's keyserver (http://keyserver.ubuntu.com/).
I don't need to run this workflow very often, as it'll take about 40 minutes to rebuild and push. But if I do update my SSG I know it'll end up in my debian repo with a version bump, so I'm happy.
On a second pipeline I can just do a simple 'add-apt-repository' and 'apt install'.
https://github.com/mauidude/go-readability
This is the kind of thing pocket/instapaper do to extract the main content from a page in a format that's easier to read (and also probably to programmatically modify)
Coincidentally, just two days ago I too went through my RSS subscriptions and, fortunately, most of them were still online, but sadly quite a few hadn't been updated in years, so it's likely that this will happen to them. Safe to say, I agree wholeheartedly with the author.
I noticed that too. I would be happy to find content directly on the archive (which is a complex search problem) and be able to download a snapshot of a website (and all its assets) from a certain point of time. I would even pay for both services.
Cool concept, although the content itself does not live in the git repo.
blog: https://gourav.io/blog repo: https://github.com/GorvGoyl/Personal-Site-Gourav.io/tree/mas...
Build with next.js, mdx, tailwind and deployed on vercel.
Even better: no tooling at all. Just files. Means HTML then. And self-contained e.g. per year.
So you don't need to touch old stuff when changing style or production tool.
Any good tool to extract a website from archive.org?
(I discuss it on my https://www.gwern.net/Search tutorial and use it every once in while eg to make my mirror of 'Climb Mount Improbable' https://www.gwern.net/docs/genetics/selection/www.mountimpro... or 'Hard Truths From Soft Cats' https://www.gwern.net/images/hardtruthsfromsoftcats.tumblr.c... )
I might still be able to get around it, but it's still a pain.
I only know of emacs markdown where you run an external process to view it (and that requires and extra install or whatever)
And it seems some themes are using this feature, therefore one can not say that Hugo is a single-binary static site generator anymore.
Just stop, use asciidoc.
For a full blog, I would use a static site generator, but not for something well-served by a single static page.