> The browser will end up rendering a tree of nodes regardless of its representation, and there's no reason that can't be inspected.
Well, wait, how much do we disagree here?
When I talk about HTML, the tree representation is the part I care about. I'm not delivering new HTML documents for every change when I write an app, I'm using something like JSX, or Hyperscript, or even in a pinch the core DOM manipulation libraries. I can still define my interface inside JS, and in the future once the bindings get set up I'll be able to define it using WASM in any language.
> I'm not sure what you mean here, can you clarify?
If an app has an accessible view and a visual view, it's possible for the features that each "version" supports to diverge. Blind users will often be given a lower quality alternative as an afterthought rather than a full-featured application.
By forcing the visual representation of the app to be based on top of the accessible (text-based) representation of the app, we can (mostly) force the two experiences to be the same. Escape hatches aside, the developer can't neglect the accessible interface, because adding features to the accessible interface is the only way to get them into the visual interface.
> Not to mention that the web doesn't really fix that problem. If the browser was poorly written and had bad HiDPI support, then so would all your web apps.
Doesn't it? It's not that the browser is scaling the web page, it's that CSS by default handles HDPI and reflow properly. Any browser that implements CSS properly will have good HDPI support, the spec is good at scaling and handling content reflow at its core.
When I call out Linux, I'm specifically calling out the idea that we needed a special Gnome mode that rendered our applications to a separate buffer, resized that buffer, and then printed it to the screen. To me, that means the GUI toolkits these apps were based on aren't being used in a way that can handle multiple font sizes, device orientations, input methods.
> I fear you may think I'm arguing in favour of something I'm not.
Maybe I do?
You bring up Swift UI, but Swift UI is a declarative, tree-based user interface that's accessible by default. Part of the reason accessibility works on iOS is because of the design patterns that I'm talking about above; design patterns that came from the web.
Swift exposes a set of common semantic components that developers are forced to use, those components have common controls built into all of them like standard ways to adjust sliders. And what limited styling options that Swift offers look a heck of lot like CSS -- just a bit more limited and managed through attributes. I'm immediately reminded of something like D3 when I look at example Swift UI code.
So when you look at Swift's declarative UI vs something like the DOM/CSS, what's the core difference that you see? What's the thing that you wish the DOM was copying from Swift?
To me, the big difference between Swift and the DOM is just the abstraction level. Swift's declarative UI is a higher-level framework that forces more visual consistency and provides more high-level tools, at the cost of being much less customizable. But one of the big ideas behind systems like React and Components is that you can write interfaces using high-level components that conform to a specific style guide -- you don't have to work directly with divs.
So if the problem just boils down to "HTML should have more high-level components like modals and popups", that's something that we try to handle in userspace, because (frankly) the web has a much wider reach and platform support than Swift does, and standardization of how everything looks and acts is a harder problem to solve on the web for high-level components than it is on iOS.
That being said, the tree-based, semantic, final interface is the part of HTML I care about. If someone thinks we can do better, and they think there are better primitives we could expose, I'm not opposed to that. I don't think HTML is perfect, and I certainly don't think CSS is perfect.
If the disagreement boils down to that, then... I dunno, maybe there's not as much disagreement then. Typically when I talk to people about this, they are arguing that we should be replacing web apps with drawing instructions. Look at the linked talk from maxharris you originally applied to -- it's explicitly talking about getting rid of a standardized, semantic tree-based text representation of the current state and handling interfaces by pushing pixel buffers.
But it kind of sounds like that's not what you want, it seems like you want a different set of HTML tags and a different document reflow model in CSS? If so, that's a position that I'm much more sympathetic towards, even if I would want to hear more details before jumping on board.