Rewriting my blog in Rust for fun and profit
jonashietala.se
jonashietala.se
It was actually good to try a few things almost blindly and then slowly narrowing with a bit more thought process. Been a long time since I had such fun investigating unknown codebase. This even helped me move my post processing scripts to templates. And then I went a step further by reading about meta tags (things like image, description, etc) and adding them via an extra variable in the post frontmatter.
Would I have picked Jekyll today? no. Is it worth the cost of switching? no.
Instead, I think the killer feature of my static site generator is going to be exactly what makes it difficult for others to use: its tight coupling with my specific use case. I'll be able to hard-code configuration knobs. Avoid dealing with edge cases. Bake strange non-generic and single-blog-post-specific logic right into the tool.
(The reason why I want to build my own is because I've tried using someone else's. And I just keep getting bitten: someone else's tool gets so big and moves so rapidly, that by the time I get around to updating my site, something about the new version of the tool has broken the way my web site works. And then I usually have to spend hours figuring out how to fix it, which is hugely deflating because I had just worked up the energy to write a new post. My naive belief is that if I have my own tool that evolves at precisely the pace I want it to, I won't run into that specific problem. I'll run into other problems of course, but I'm thinking that I'll prefer those problems to the ones I currently face.)
I have been looking to break out parts of static-site-generators to make it easier to bring one up. So far I only have two primitive building blocks
- https://docs.rs/engarde/latest/engarde/ - https://docs.rs/file-serve/latest/file_serve/
I can't even imagine I need a template system beyond something I can write in a few dozen lines with some regexes.
I did the SQLite thing more than a decade ago. I ain't going back to it unless I have a really really good reason for it.
I really just need a purely static web site where I can write Markdown posts. There's some other stuff I want in terms of producing the static content, but in terms of web site functionality itself, it should just be HTML, CSS and some images.
I also don't mind re-compiling. For my needs, "compiling" my web site is likely to take less than 1 second.
Definitely don't need plugins. :)
I meant it. I'm done with static site generators built by others. They have caused me too much pain.
EDIT: That "oops" is already reality: https://github.com/Fizzadar/Luapress#luapress-v4
Static site generators are just too easy to build and too easy to let die. And the ones that don't die grow big and bloated. That's my experience. The only remedy I see is to build my own and simplify ruthlessly through tight coupling.
"You either live long enough to see yourself become the villain, or you die a hero."
Sadly you're right. I've been monitoring in the trenches for a while and what you said is covering it perfectly.
RE: maintenance, I was aware but was wondering whether Luapress is mature and complete enough to not need maintenance. I'll give it a go.
And I truly get your point about point releases of languages breaking stuff. I am severely burned out on that front myself that's why my next game plan is to buy and setup a Linux workstation and just isolate various software pieces that I [might] need in Docker containers and periodically backup those for extra good measure.
I am done with all that constantly moving and breaking crap as much as you are. Just looking to find software for all my needs and nail its versions and installs and just allow myself to not constantly keep up with the youngsters' play toys.
Appreciate all your software contributions and a good part of them are my favorite in their area.
--
For Nix, I think you know what I'll tell you: you don't recommend a war veteran with PTSD "just one last war". I get what they're trying to do but their discourse on various platforms has left much to be desired and they seem opposed to offer more ergonomic CLIs and/or UIs. I hear that's changing as well so I'll be checking them out a few times a year. As it is right now, I wouldn't even mind the steep learning curve -- I am not completely burned out and I still punch quite hard in my work -- but mysterious error messages and maintainers ending discussions with "well then it's not for you" et. al. are just not appealing and feel like I'm taking a risk that will not pay off. Nix still feels like somebody's experimentation project.
I truly hope they take off as their idea deserves but that must come with simplicity on the level of your average Joe and Jane pasting 2-3 commands in a terminal (sort of how you install Rust; it's literally two pasted commands). Before something similar happens, Nix is doomed to remain a niche curiosity.
My personal website is served by a couple hundred lines of Rust. There is no database or ORM -- I just defined structs like `BlogPost` or whatever and stuck a `#[derive(Serialize, Deserialize)]` on it and use the `serde` library to JSONify it, store it flat in a hard-coded folder.
I even wrote my own dead-simple Markdown-ish flavor that is tailored exactly to my writing style. I use a lot of em-dashes, for example, so it translates `--` into `—` (rather, &mdash). Or, since a lot of my posts link to other posts, I can do `[post title_of_other_post]` and it'll generate the appropriate link.
Whenever I want a new feature while writing a particular blog post, e.g. pretty inline SVG, I just go add it and then keep writing.
This is true, as a general principle, not just of static site generators, but of all software, and is one of the major Reasons Why We Can't Have Nice Things, and why automatic updates are evil and poisonous, and manual updates dangerous and often-avoided. See eg https://jacquesmattheij.com/why-johnny-wont-upgrade/.
I don't think so. There's lots of software that I've used and upgraded for decades and rarely if never run into any issues.
I don't think it's a general principle. I don't like generalizing to "all" unless I have a really good reason to do so. Instead, I stuck to my specific experience with static site generators.
I guess writing your own non-generic own-static generator should be relatively easy and worth the effort as you wouldn't have to read some obscure documentation (like Jekyll).
I've since taken up photography, so now I need to automatically read metadata from jpegs, add some tags and organise it on the site in albums. It's a bit outside the reach of XSLT, so first I've tried using Python with Jinja2 and PIL. Same problem as previously. So I've rewritten it in Go with the help of encoding/json (for the manifest), html/template (for the template), image/jpeg, golang.org/x/image/draw (for generating thumbnails of different sizes for different screen sizes) packages and exiftool (reading and removing EXIF tags). Pandoc has been present throughout all the rewrites, because I like its flavour of Markdown and options to shift heading levels by a specified value, so that they fit in with the surrounding template. I'm putting some finishing touches and am going to upload it in the coming days.
I know, you stopped reading at "XSLT". :D
One of the most annoying aspects of working with hugo was the friction from idea to publishing it.
I decided to go a completely different route and built https://prose.sh
With prose all I have to do is to create a markdown file and then `scp` it to prose.sh. The platform takes care of the rest and I can still enjoy my own editor and terminal tooling. Even better, authentication happens via SSH so the friction is extremely low.
It's a nice balance between writing a blog on a platform that does everything for you and being able to use my own tooling to write content.
Another big issue with Hugo is when I want to make edits to the content. My brain is weird and I only scrutinize my content after it has been published. So I end up making a lot of edits after an article has been live for a few minutes. It's a terrible habit and Hugo required deployment for every little change.
I've thought about converting prose to be SSG -- we could generate the html when a user uploads the file -- but honestly the site is fast as it is so I don't see the value in that added complexity.
I always set mine to 2.
How do you handle images? That’s my main annoyance / biggest hurdle when publishing. I have to take each image, resize it for web, put it in static folder, and reference that pathname in the MD file. Not a big deal, but would love an easier way. I suppose I could write a simple shell script that does all of the resizing and prints out a path back for me to use.
Presumably there are reasons people want to use more specialised tools, but I don't have any of those.
I made a few attempts in the past at building more general-purpose tools that would cover my use cases, but I've never followed through on them. Too much of the work was in the bits that don't actually benifit me right now.
My new tool is simply called "Other Jeff". I write onto it what I need right now. Do I have a series of pull requests that need babysitting to get them over the line? Other Jeff learns to poll their status and notify me of important events along the way. If I want to store some state about what I'm working on right now, I need a database of some kind, right? Nope. Not any more. I hard-code it into `db.rs`, because it turns out there's almost nothing I need to persist that's too burdensome to update by hand.
So what I have in the end is a single Rust program that grows and shrinks depending on my needs at the time, accumulates reusable helpers if and when they actually make sense for _me_, and slowly learns to automate more of my daily distractions.
But I appreciate the link to Zola. I haven't looked in this space for awhile and it looks good.
One issue I hit is that taxonomies aren't available at the section-level (nor are they section-specific, in case you want to treat sections as sub-sites), this makes it seemingly impossible to have a section taxonomy view.
I might try to patch it, but it's tricky as it's like merging the taxonomy code with the page and section rendering.
If I were to make my own generator from scratch, I'd go with python. It's the language I work with the most. A lot of my blog content will be generated from python (figures etc) or will show python code. So having a site generator in the same language can help me add features by hacking on the generator code in a way that a single binary cannot.
It's tricky. You can't accept all the features otherwise you just have a mess. At the same time you need to also add/supports features that a lot of people want but _you_ personally don't need if you want the tool to be used by more than 1 person. For example with SSG I don't use themes, I just write my own templates for every site. Despite that I still added it to Zola because most people want themes. Some people will not be happy when you say no to some features they want but hey, it's open-source they can always fork it.
I'm considering using Pandoc with Soupault to my website markup agnostic by being dependent on Pandoc. Soupault can act as a HTML processor although I'm not sure if that's enough to not need a template langauge. Or maybe I'm mistaken about Soupault.
That's it.
Works fine.
Not everything is improved by creeping featuritis.
I still haven't released it as there are some bugs (localization!) that i need to fix but if you want to have a play the staging site is here[1].
Edit: and... just read his bio
And of course don't write a plugin system or something you would share with others! Just write the thing that you need, and change it as you need to!
Static site generators are great for people less comfortable doing this, but having to deal with all the _stuff_ is a major waste of time for what it's offering for people who don't need it.
The best path depends on your personality, but building my own blog tooling killed my writing. I'd start writing something only to realize I needed to tweak something. Then I wouldn't even write the post or I wouldn't have the energy to fix the tooling and I'd do neither. And then I stopped writing.
I didn't start writing again until I used Wordpress which let me focus on writing, though after proving myself for a year, I switched to Jekyll + Github Pages so that everything was on Github.
The one time static stie generators are helpful is if you really, really like the out-of-the-box look and behavior of some existing theme.
I think it's tempting to add a bunch of "time saving" features for writing, but lots of time just going with the HTML is fine. Maybe if you're writing out a post every couple of days then you can focus on time saves. Until then MD will get you most of the way, and otherwise you can always hack in some special case for your unique post
<aside>
Markdown syntax is *not* interpreted here, and this is a text node inside an aside.
</aside>
—⁂— <aside>
<p>Markdown syntax is *not* interpreted here, and this is a text node inside a p inside an aside.</p>
</aside>
—⁂— <aside>
Markdown syntax *is* interpreted here, and `aside > p > em` should match.
</aside>
—⁂— <aside>
<p>Markdown syntax is *not* interpreted here, and this is a text node inside a p inside an aside.</p>
</aside>
—⁂— <aside>
<p>This line is a code block (pre) inside an aside.</p>
</aside>
—⁂—And all that is without taking into account variations in Markdown flavours, and what may happen with unprincipled and inconsistent extension of Markdown. All up, I say if you want sanity that will last, go HTML and avoid Markdown strenuously for anything beyond simple prose.
Seriously though, I think that what you're describing is a real failing of Markdown. I wanted at one point to use Restructured Text but really you want Sphinx and Sphinx is not meant to be used as a library unfortunately.
I have a well-established track record of disliking Markdown because of how grossly technically unsound it is. I’m steadily designing a more principled replacement for at least my own use, drawing most significantly from reStructuredText, AsciiDoc and Markdown.
I'm working on trying to write a static site generator for my own personal site/blog in C++, been tinkering for a couple months. It started as a passion project in tribute to my dog that passed. GatsbyJS wasn't working (again) which is what my current site is built with, so I just said screw it, I'll write my own. Chose C to begin with and quickly gave up. Decided to switch to C++ because its what I'm supposed to be learning for work.
I named it bluesky, after my dog that passed, Sky Blue. https://github.com/mas-4/bluesky
Templating is a lot harder than I had initially thought. I wanted the templating system to be as bare bones as possible, similar to another static site generator I rather like, sergey[0]. Most of these articles I found ([1], [2]) about writing your own SSG don't go into templating much, they just use an off the shelf library. Inja[3] is available for C++ but, like I said, I want something really bare bones, like if you were designing html now, you'd obviously include html-includes and html-templates. I finally got it working for includes and templates, now I have to add the markdown support, and then I plan on migrating my personal site to using it.
[0]: https://sergey.cool/ [1]: https://www.smashingmagazine.com/2020/09/stack-custom-made-s... [2]: https://blog.hamaluik.ca/posts/build-your-own-static-site-ge... [3]: https://github.com/pantor/inja
edit: a link
Templating languages probably help with performance a bit when you're rendering at request-time, but for static stuff I don't really see much point
I've got a hand-rolled static site generator in JavaScript that just uses template strings. I only have around 30 pages, to be fair, but when I make a change it re-generates the whole site before I can even alt-tab and reload the page. I assume a C++ version could be much faster
In a server-rendered site you have to worry about injection attacks, but that's not really relevant when you're statically generating a site from content that's fully under your control
It is very convenient to have an automatic escaping of everything.
The problem is that it does several pandoc parses for each page:
1. Plain text - Input for audio parsing (which takes a significant amount of time, I have a version of the build that skips this).
2. Simple HTML - Input for RSS feed.
3. Full HTML - Static web page.
One change I would make to pandoc is to have the option for several outputs (as was discussed years ago [1]). I think generating multiple outputs is highly common usage.
One feature I will get around to writing (eventually) is automatic image compression. I had some success using jpegoptim and optipng on another project. This takes a while, so I want to build a command line tool that caches the output of a command if both the input file(s) and command do not change, and invalidate based on time. That way I could blindly do something like:
image="cat-pic.png"
cache-run -i $image -c "optipng -dir out/ -o 4 -strip all -fix $image"
(You might want to experiment with different optimization values based on resolution, etc).[1] https://groups.google.com/g/pandoc-discuss/c/lex900rSpOM
I’m a huge fan of treating one’s blog as a “homelab” for exploring new (or old) software stacks for the web, so this idea resonates a lot.
One thing I’d like to try (has anyone tried this?): eschew the idea of a static site generator and go the “old school” route of dynamically generating each page on the fly, but take advantage of the modern CDN layer of the web (CloudFlare, etc.) to effectively cache and give you the speed equivalent of a static site. Yes this approach has more dependencies but I think simplifies the work done on the blog side since you can build it using whichever simple web framework appeals to you most (or none at all…this is a homelab after all).
My plan (don’t we all have this plan) is to try this out in Crystal when time allows…
I think the best thing about this approach (writing your own generator) is that it really lets you build your site / content exactly the way you want to. Don't need categories and tags? Don't add them to begin with. Want to add an RSS feed? No problem, just add it. Want to do some fancy image compression and processing in the pipeline? No problem; you're the boss. It sounds simple but it really frees you up to write / generate your site however you want.
> It’s annoying to install all these and to keep them up-to-date, it would be great to just have a single thing to worry about.
I feel the author's pain.
There's a Python2 project for ebook management that I used pretty extensively years ago, but it depended on a number of external dependencies, both as compiled shared objects and other Python libraries. It felt extremely fragile and was always a pain to get it working properly, so I eventually rewrote the core parts of the entire project for my own use (in Rust, incidentally). Now I have a single app that I can be sure will pretty much always work.
Though Calibre has at least made it to Python 3 these days.
Eventually I really appreciated the Golang approach to templating much more than Jinja/Liquid style templates, Go has kind of a composition-over-inheritance approach to templates.
If I ever have time to get proficient enough with Rust I'd love to port that Go templating approach to Rust. Maybe not so much the data interpolation syntax, but rather the approach to composing template partials. But ultimately I'm familiar enough with Jinja style, Zola would probably be a breeze to adopt as well.
It cites concrete examples of where the rubber really meets the road: e.g. Cargo as a one-stop-shop for platform insensitive, “just works” build and dep management.
It acknowledges that it’s ok for fun and novelty to be features, which it is!
It steers well clear of A >> B nonsense. No preachy TED talk about memory safety.
It’s respectful of the value that other languages provide in their respective niches.
This is how to sell Rust. It’s novel and fun but also pragmatic and industrial grade! And if you’re into X, that’s cool too!
Here are some select quotes from the article:
> Please note that this is not to say that Rust this [sic] much faster than Haskell, instead see it as a comparison of two implementations. I’m sure someone could make it super fast in Haskell.
> I don’t think it means that Rust is simpler or easier than Haskell, it just means that I personally had an easier time to grok Rust than Haskell.
I don’t know how Rust, which is quite a cool language, attracted this sizable bloc of aggressively intermediate jihadis, but it’s really worrying because Rust is in fact innovative and important, but epic languages have crashed and burned over less.
And then you’re going to act like you’re the self appointed arbiter of discourse on HN, saying that this kind of article is ok but that kind isn’t? Maybe a person who calls those they disagree with jihadis isn’t qualified to police what others say.
Let’s talk about the way you talk. Here’s you from 5 days ago accusing a good tool of being a “dumpster fire” and saying “it crashes process IDs more often than Justin Bieber crashes Maseratis”. (https://news.ycombinator.com/item?id=32581005). But it gets worse, because later you start crying about being downvoted.
At this point it looks like a speed run of how many HN rules you can break. And then you follow that up with tone policing others.
- Why Not Rust? (https://matklad.github.io/2020/09/20/why-not-rust.html)
- A Rust match made in hell (https://fasterthanli.me/articles/a-rust-match-made-in-hell)
Compare these with the superficial criticism that benreesman was making. He couldn’t even link to the proper issue, a fact he realised 10 hours later. Even had he linked to the right issue, he has no idea what he’s talking about. He says he had some issue with the language server and then simultaneously mentions LLVM? rust-analyser doesn’t import LLVM because it doesn’t need to generate machine code. You think when such poor feedback gets pushback it means people aren’t listening. No, it’s just poor feedback. If you want high quality criticism of Rust, read the articles I linked. Rust has many shortcomings and these folks discuss it in detail.
> no other language has a rewrite it in rust meme.
Your account is a year old. I’ve been here almost 10. I remember when everyone was excited about Go and wanted to rewrite everything in Go. And before Go it was Python. A couple of years from now they’ll be rewriting everything in Carbon.
This is normal when things are new and people are still discovering it. Most of the folks you’re encountering are excited by Rust … and that’s ok. It’s not a bad thing to be excited or share that with others. If that reduces your quality of life, I apologise. But they meant no harm.
> No other language
Firstly, no language does anything. And secondly, you’re making sweeping generalisations of a large community based on a couple of incidents. I don’t know there’s anything I can say, other than I haven’t seen it in years. No one thinks it’s a good idea to attack anyone, and moderators are usually quick to nip any anti-social behaviour in the bud.
Lastly, anyone with a lick of Rust experience knows that unsafe is useful and widely used. You can’t implement std::Vec without it, for instance.
None of those languages had as many rewrites as rust (which is a testament to how good the language actually is) nor did they have the same eventually obnoxious meme.
And yes, I am making sweeping generalizations about the community based on a couple of incidents. Incidents that have never happened in any other language communities I've participated in. Didn't official rust moderation team essentially quit relaticely recently because some business with the core team? I don't really care about the details. Rust is too smart for me, but I like using it anyway. I'm happy to stay separate from the community around it though.
Maybe you just have an axe to grind?
Dude, touch some grass
HN became boring in the last few years anyhow.
On the one hand, it’s depressing that Cargo is forced to “go it’s own way” in order to achieve some sanity because it can’t can’t trust the platforms, but working is working and I can’t remember more than one or two times I’ve seen it shit the bed. It’s pragmatic, useable, discoverable, and industrial strength.
I've used Rust and Cargo a bit but I don't have a lot of insight into it, what does Cargo actually need to implement itself instead of trusting the platform?
Distros do still package Rust, and some are even relatively up to date, but finding third-party libraries that work with [insert arbitrary old version of Rust] can be a challenge, because libraries often start using new APIs quite soon after they hit stable (and obviously old versions of Rust don't know what to do with that; this is the difference between backward compatibility and forward compatibility).
It isn't as if it wasn't done before.
I agree with a lot of the points about Rust too. Every language is a selection of tradeoffs and for me Rust's chosen tradeoffs I think are both important (eg memory safety) and reasonable.
Take something like Cargo. The author says "It just works". This is my experience as well. But more importantly, dependency management and packaging were built into the language and ecosystem from the start. It still boggles my mind that any language that came after Java didn't do this (yes, I'm looking at you, Go).
My heart did sink a little bit at "Using Regex of course!" though.
Last time I did it was with Elixir + Phoenix + PostgreSQL, without any standard or reasonable approach. It was a lot of fun and still using it today.
https://www.brightball.com/articles/insanity-with-elixir-pho...
But since I have done that, I feel that many existing site generators are bloated, probably, because they cover scenarios, which I would not want on my blog anyway.
347 options and counting!
> But I haven’t used Haskell in years, and the overhead for me to add slightly more complex things to the site is quite large
I’m starting to come around to the idea that it might be better to avoid shiny fun things in favor of having most of my active projects use the same stack, or at least the same few languages.
I’ve got a blog in Jekyll (with some Ruby-based customizations), a bunch of stuff in JavaScript, some TypeScript, an app in Elixir & Phoenix, a (desktop) app in Rust, and some Rails to maintain. For me personally, Elixir is the odd one out: I learned enough to build that app, and I really enjoyed Elixir, but that knowledge was fleeting because I haven’t needed to touch the code very much. I have to relearn how Phoenix works in order to make changes. As much as it pains me to say this, I kind of wish it was in TypeScript or something, just to reduce the mental overhead. (but then, that’s a whole other can of worms because of how piecemeal things are in the Node ecosystem, so I’d end up having to learn a web framework and an ORM and some other libraries, which feels like the sort of knowledge that might be even more fleeting!)
If you once knew how Elixir works (syntactically), then you will be able to quickly bounce back. If comparing TypeScript with Phoenix, then you would have to relearn your in TS written functionality, in contrast to relearning Phoenix.
I don't see how it being TS would help much tbh. I think what matters much more is, whether your past self left good comments and wrote readable code.
Knowing a lot of languages well is a super-power in some ways, but heavily polyglot environments create real pragmatic challenges. There’s a reason so many marquee shops carefully curate a (short) menu of production languages.
It’s been my experience that if (and it’s a big if) you have either the need or inclination to operate a big orgy of polyglot language interop, the hellacious learning curves on NixOS and Bazel are “The Way”. It’s a big investment and only worth it if you’re pretty committed/obligated, but there is a path for those whose needs justify the cost.
Well, email is a actually a single file CMS which contains html, plaintext, attachments/images, and metadata like when the email was sent and the subject line,
With all this, it turns out you can create pretty nice looking webpages without having to futz around with md->html conversions.
The huge advantage is that since I know gmail, whenever I have an idea for a blog post, I can go from idea to published in minutes. The biggest time suck tends to be finding stock photos for the post.
As far as tech stack goes, it’s all serverless. I use SES to handle inbound emails and the excellent nodemailer package to parse .eml files.
Wow, what on earth was Hakyll doing all that time?
Because Rust binaries are so secure, you can expose them over the internet safely. With a little bit of glue, you can automatically recompile the CGI binary every time the rs source is touched, for a dev experience similar to interpreted web scripts.
That doesn’t stand to reason. Safely in that you can confidently avoid things like remote code execution, sure. But that doesn’t necessarily protect you from denial of service attacks; for example, I expect most services built on Hyper to be vulnerable to slowloris attacks <https://en.wikipedia.org/wiki/Slowloris_(computer_security)>, because I believe it still defaults to basically no sane kind of timeout, and doesn’t provide the tools to handle it properly anyway <https://github.com/hyperium/hyper/issues/1628>. If that’s part of your threat model, you should avoid exposing it directly and use a reverse proxy.
Good somewhat related reading, on a technologically simple DoS attack on Rust-based code that was largely fixed by speeding things up, fixing caching and proper use of reverse proxying (via Cloudflare because it was genuinely heavy load): https://fasterthanli.me/articles/i-won-free-load-testing
You can of course still do stupid things inside the rust application that explode your number of concurrent requests, like waiting for slow external resources, databases etc. For a simple static formatter, it should not be the case, it would just repeatedly hit the disk cache and some minimal CPU.
... but with an execution speed that consumes negligible CPU time from the moment the template/source is loaded into memory until it is formatted and put on the wire. Compared to that, a PHP template is dog slow. I don't have any direct experience with Go, small GC allocations for string handling might introduce non-negligible overhead.
Have the tenets of Make (and the like) from over 40 years ago been forgotten?
I guess maintaining a dependency graph just isn't a high priority feature for site generators.
On that topic, I wish there was a generic build system like SCons written in something faster. Maybe there is (does Bazel fit my bill?) and I just don't know of it yet. What I love best about SCons is that you can just write custom builders right there. No need to drop in a shell script - you have the much better Python standard library at your fingertips. This also means it's way easier to make cross-platform.
Unless you're 'web-scale', it's complexity that doesn't solve any particular business need.