> You need at least "real world units" for stuff like the minimum size of touch targets [...].
That touch targets need absolute lengths is a good point.
> So that makes for three independent settings, in the general case.
Yes, and you’ll notice that’s what Darktable does (or did) :) I’m not unsympathetic, just have never seen it done in a moderately complex situation.
> Modern CSS is quite close to Turing complete, so an entirely constraint-driven design (which is what's needed when more than one unit is being chosen by the user) ought to be feasible.
I suspect it’s not a question of theoretical expressive power so much as ergonomics: allowing easy description of (what people think about as) simple things. It probably doesn’t have to be precisely; if a full constraint-based / linear programming approach[1] is the way to go, I expect designers will adapt, it’s just that last I checked those could be quite tricky to use: small changes could lead to drastic rearrangements of the layout, and as soon as you try to add a couple of breakpoints the whole thing becomes NP-complete.
It’s like, say, LR(k) parsers. LR(k) parsers are nice, fast, and good at expressing our intuitions about ambiguity (unlike ordered choice). They are expressive enough for almost anything you might want to say. But nobody really wants to work with them, the failures are tricky and hard to predict, and they lack many things you want for modularity (say, closure under unions).
It seems to me that the situation with truly resolution-independent layouts is similar in this respect: we can technically do them, to some extent, but they are a pain, so nobody does. And it’s not so much a problem of finding the perfect language as a problem of figuring out what it is that we want to say that seems so simple in our heads.
[1] https://gss.github.io/guides/ccss