A Clean Start for the Web
macwright.com
macwright.com
Structural, not societal.
And Google's monopoly is not good, but it's not really the root cause of the issues articulated in the article.
Linux, Wikipedia and a number of other proves this is not the entire story.
If the new <x> is good/fast/cheap enough people will sometimes start using it.
Often this is gets great help from the incumbent solution being either really bad and/or slow and/or expensive.
This way of thinking is not just not entirely correct but more importantly it is demotivating. Maybe what we do won't succeed but for me it sure beats watching more TV :-)
I agree on the second point, but it's more by accident in my opinion. If the technical solution leads to people finding it easier or cheaper to do something, they'll adopt it. Whether it's new, good, pristine, perfect, free & open etc, doesn't matter. It has to make their life easier, save them time and/or money or otherwise enhance the experience, that's what counts.
But new technology has had a dramatic impact on societies in the past, disrupting monopolies, creating new ones, confronting society with new choices, opportunities and challenges.
Technology can force a rethink. It cannot force outcomes.
Let the different subsets compete. You'd have multiple different simple browser that each support a different subset. Developers would have to choose which one they agree with and make their sites compatible. They would all render properly in full browsers. You can find the sweet spot over time.
It's way better to have one jankier universal platform with a few good implementations.
Yes, people have gravitated to centralized platforms for various reasons, and those platforms tend to make it difficult to publish entirely handwritten raw HTML/CSS, but it's not because the alternatives ceased to exist.
Since the XML could carry pretty much any sort of data server endpoints could back an "Application Web" just as easily as a "Document Web". A browser could take a music playlist and transform it into a listing of tracks via XSLT while a music player would take that exact same URL and enqueue the tracks for playback.
Unfortunately the mixed presentation and content model had a lot of momentum. By the time a lot of web developers caught on to separating concerns XmlHttpRequest existed and let everyone reimplement all those cool XML/Web 2.0 ideas poorly in JavaScript.
Honestly if I actually improved my tools a bit better, I'd almost enjoy using it. But it sucks, so damn much, in the current state.. And now nobody seems to be interested in XSLT ("XML is legacy"), I don't think I'm ever going to be interested in "fixing" it.
Its regrettable that such a vast ecosystem exists for a markup language to facilitate its use in data transfer, transformation and validation.
It seems like a waste that so many working standards seem to have just been tossed away with nothing to put into their place because Json saves a few bits on the wire.
Also regrettable is the fact that there seems to be resistance in developing similar standards for json.
With a standard, however, I can be sure that what I intend is being validated and if my logic is wrong then I can change it to match the spec.
I find JSON to be a lackluster replacement for XML. The efforts to bolt on XML features (JSON Schema etc.) are half assed and not widely supported. Without a bunch of extra complication you can't really know what a value in JSON means beyond its translation to some primitive.
The web development world doesn't deserve XML! /s
I am a bit of an XML fanboy as it solves complicated problems in a logical way. I also bought into (and continue to believe) in a lot of Web 2.0 ideas beyond pastel colors and pseudo-reflections on logos.
1. Use HTML and vanilla JS or jquery without using the huge set of javascript libraries needed to render a simple page.. as an SPA. What was the problem with plain old HTML in the first place? and whats the problem with server side rendering?
2. Avoid Google Adwords and the rest of the nonsense.. including AMP.
3. Support search engines like DuckDuckGo
Now, none of this means your website will be visited often or find its content listed high on Google.
However, if you can live with that, why not?
Split the Web? On the server side? Or the client side?
Splitting the server seems unnecessary - since servers already let you serve as simple a site as you like. If you want to make simple sites with no javascript you are welcome to do that.
So split the client? Have one app to read some sites and another to read some other sites? That sounds contrarian to literally every program evolution in history. Because all client software strives to read all relevant formats (think Word loading WordPerfect files.) why would a user choose 2 separate apps when 1 would do?
It seems to me that the author misses the point that sites are javascript heavy not because they have to be, but because their makers want them to be (largely so they can make money)
It's definitely not easy by any definition. And also not happening accidentally. Almost all frameworks/libraries mention about how small they are in overhead.
Otherwise, the problems he s describing have 100% to do with the miserable culture of CV-padding that has taken over the silicon valley and led to SPA monstrosities, rather than with web standards themselves. WWW as a standard is OK, old websites work, and HTML/CSS is still there. Trying to make the DOM work as some kind of window rendering engine is a travesty that doesnt work.
Regarding the CV padding theory the explanation is more straights for me: it simply works better that what we had previously for enterprise software. It’s more relevant to compare SPA not to the WWW but to 90’s or 00’s corporate software.
Like what would you use for the typical CRUD with some analytics business app? The drawbacks of SPA are nothing compared to the universality of the web interface (“oh and it must work on my iPad too”) in the enterprise world.
I don’t understand what you’re saying here. SPAs aren’t circumventing web standards, like Flash — web standards are what make SPAs as they currently exist possible in the first place.
Indeed, the complaint is that the standards themselves have gotten too complex, such that only Google —with its obvious conflicts of interest — has the resources to develop a browser that complies with them.
I’ve experienced this. I started web development in 1996/1997, so I have a lot of experience building pages by hand. These days I do use Vue.js, but sometimes I’m still known to build pages by hand. I’ve gotten so much grief for it over the years. Not because there is anything technically wrong with my files, but because “it’s not the way things are done today”.
> You might want to start with a lightweight markup language, which will ironically be geared toward generating HTML. Markdown’s strict specified variation, Commonmark, seems like a pretty decent choice.
Markdown of any form would be a catastrophically bad choice for something like this. Remember this: Markdown is all about HTML. Markdown has serious limitations and is very poorly designed in places, but it works well enough despite being so bad specifically because underneath it’s all HTML, so you can fall down to that when necessary. If you go with Markdown, you’re not actually supporting “Markdown” as your language, but HTML. If the user writes HTML in their Markdown, what are you going to do about it? You’re supposed to be able to write the HTML equivalent of any given Markdown construct and have it work, and I find that perfectly normal documents need to write raw HTML regularly, because Markdown is so terribly limited—so your document web is either (a) not actually Markdown and uselessly crippled (seriously, plain Markdown without HTML is extremely limiting for documents), or (a) actually completely back to being a subset of HTML, which is exactly what you said you didn’t want it to be (rule #1).
If you want to get away from HTML, you want something like reStructuredText, where you don’t have access to HTML (unless the platform decides to set up a directive or role for raw HTML), but have a whole lot more semantic functionality, functionality that had to be included because HTML wasn’t there to fall back on. I think AsciiDoc is like this too.
But the whole premise that “HTML must be replaced because it’s a performance or accessibility bottleneck for documents” is just not in the slightest bit true. HTML is fine. Even CSS is well enough (though most pages have terribly done stylesheets that are far heavier than they should be, e.g. because they include four resets and duplicate everything as well as overriding font families five times). The problem lies almost entirely in what JavaScript makes possible—not that JavaScript is inherently a problem, but it makes problems possible. The proposed variety of document web would be unlikely to perform substantially better than the existing web with JavaScript disabled.
All up, Robert O’Callahan’s response is entirely correct. You can’t split the web like this; you will fail. Especially: if your new system is not compatible with the old (which your rule #2 precluded), it will fail, unless it is extremely more compelling for everyone (rule #3), but I don’t believe there’s anywhere near enough scope for improvement to succeed.
Originally, this was the case, but I'd argue that modern MarkDown flavors have been separated from HTML. It's common to compile via other pathways than HTML now (e.g. MarkDown → LaTeX → PDF), and many implementations don't support inline HTML anymore (e.g. many MarkDown-based note-taking apps).
> I find that perfectly normal documents need to write raw HTML regularly, because Markdown is so terribly limited
Out of curiosity, what do you need it for? I write a lot of MarkDown, but haven't felt any need for writing raw HTML. Modern versions of MarkDown (e.g. the Pandoc variant) have native support for things like tables, LaTeX equations, syntax-highlighted code, even bibliography management.
1. Adding classes to elements, for styling; admittedly this may be inapplicable to some visions of a document web, if you can’t write stylesheets.
2. Images. If you’re dealing with known images, you should always set the width and height attributes on the <img> tag, so that the page need not reflow as the image loads. Markdown’s image syntax () doesn’t cover that. (Perhaps an app could load the image and fill out the width and height as part of its Markdown-to-HTML conversion, but I haven’t encountered any that do this.)
3. Tables. CommonMark doesn’t include tables, and even dialects that do support tables are consistently insufficient so that I have to write HTML. For example: I often want the first column to be a heading; but I don’t think any Markdown table syntaxes allow you to get <th> instead of <td> for the first cell of each row.
1. A markup language taking multi-lingual considerations to heart is a must for me. Markdown doesn't work well for non-alphabetic languages in terms of typesetting etc. The web is global and multi-lingual, a new web (if there is one) should be, too.
2. I prefer a new content web rather than a new document web. When I use the web, I absorb and sometimes create content beyond documents. I don't want art communities, for example, to be excluded from the new web. Although an art sharing community can be modeled after documents technically, 'content' is the more apropos. Your mileage may vary, though.
Which features, in your opinion, are required for decent non-alphabetic language compatibility on top of plain Unicode text?
You may want to read the work by W3C[0]. Some requirements mentioned exceed the needs of simple written communication, but not all in my experience. The problems usually arise when you mix different scripts.
[0] https://www.w3.org/TR/typography/
EDIT: clarifications.
[0]: https://regisphilibert.com/blog/2018/08/hugo-multilingual-pa...
As in, no more of the stuff on this list [1] without really broad buy-in. We're seeing Google do "expand, engulf, devour" to the Web.
Already, it's hard to send mail unless you're a big player. Writing a browser is a huge job. Just fetching an HTTP file is now complicated enough that people run "curl" as a subprocess to do it.
* declaring JSON (or similar format, but why reinvent the wheel) variables
* rendering the values of these variables
* interacting with the variables with controls (inputs, check boxes, selects, etc)
* conditional rendering based on those variables
* iteration over them
* submitting a form as a JSON object
* saving a JSON response from the server into a variable
* re-rendering the effects of a variable change without a full page reload
I think a major reason for why we ended up here[2] is because browsers did not add these features to HTML and thus devs resorted to AJAX and small JS snippets at first and then kept expanding the use of JS.
[1] as a Java dev who used to do front-end before the ascendance of client side SPA frameworks I'm thinking of JSP, JSF, Thymeleaf, etc. Maybe the client side stuff used nowadays is similar, I'm just not familiar with it.
[2] being forced to run a turing-complete programming language just to display a fully static document, leaving one exposed to security risks (yes, the browser guys are doing a much better job sandboxing than Flash, Java Applets, etc) and tracking.
That's my hunch as well. If HTML had built-in features such as modals (maybe as an evolution of popups), or partial rendering of server side data (maybe as an evolution of frames/iframes), it would have taken more time to reach the current state.
Btw, your list reminds me of Intercooler.js.
I don’t believe Chrome’s user stylesheets even supported anything like Firefox’s @-moz-document for scoping the styles to an origin or similar—it was a completely blunt instrument that had very limited use, even less than Firefox’s userContent.css.
I wouldn’t complain if userContent.css were removed from Firefox for the same reasons (though I would certainly complain if they removed userChrome.css, since no alternative exists in that case). I used to use userContent.css. I’ve used extensions (now Stylus) for the equivalent functionality for years now.
The tweet this article cites about its removal six years ago is terrible: it quotes the most recent comment on that issue as though it were fact stated by the developers, when it is in fact a random person (probably a disgruntled user) jeering; said comment is completely false.
It's more of a spectrum than a binary, a lot of things exist at various points in between "completely document" and "completely app". Even wikipedia has editing interfaces which are arguably "app" right? I guess it could be split into two separate "things", one using the "document web" for display and another using the "application web" for editing. This doesn't seem all that appealing.
I think the ability to be wherever you want on the document/application spectrum, and have a site (or a part of a site) move along the spectrum over times without having to "rewrite in a new technology" is part of what leads to success of the current web as it is.
The OP is probably not wrong that it's also part of what leads to the technology mess. But I don't think it's realistic to think you can bifurcate the architecture for two ideal use cases like that.
Do you also believe that there can be no clean division into "data" and "code"?
If a program I wrote is interacting with an API, is it not completely unambiguous for me to ask the maintainer of the API to give me some way to directly read all of the persistent data behind the API?
I don't think it's a general principle about "data" and "code" in general", although of course data and code can also slide into each other in many other venues.
But of course you can write programs that access APIs. I'm sorry if I was unclear and you thought I was saying it was impossible to write programs that access APIs to get persistent data.
A major reason people are concerned about the recent Mozilla layoffs is because they're one of the few competing browser vendors. Why does that matter? Because browsers are too complicated for reasonably sized teams to implement competing solutions. If a larger percentage of web content was accessible to simpler software, this would be much less of an issue.
It's not hard to make simple document browsers. My personal content is accessible in the console as plain text:
curl https://apitman.com/txt/19
nc apitman.com 2052 <<< /txt/19
I like the idea of splitting the web between documents and apps (and using separate browsers to access them), but don't necessarily agree making it incompatible is the way to go. I think specifying a subset of HTML/CSS and allowing the "new web" to grow over time might see more adoption better in the long run.TL;DR: I’m abandoning HTML and switching to PDF.
I can imagine why you might have concluded as you have, but your conclusion is nonetheless baffling. Most of the reasons you’ve chosen PDF over HTML apply just as truly to HTML; page-orientation is the only one that doesn’t, and I refute your claims that that page orientation is a good thing for the web at large or the hardware that most use. (Guess how I read the document? By scrolling, not by pagination; the footer and header of each page transition is just a minor annoyance in the middle of a paragraph of text.) I disagree with every point of your assessment on PDF’s historical disadvantages, too: PDF as used still includes patent-encumbered things, and implementations are complex enough that major deviations and incompatibilities are common; PDF files are still ginormous (nothing has changed about this ever), and tooling is terrible (perhaps even non-existent?) for trying to shrink files without butchering the lot, with yours as a good example of not being an order of magnitude smaller as claimed; it’s still far harder to make PDFs accessible, because most free tools just can’t do it, and it takes far more effort than HTML where it’s easy (as before, tooling is terrible, where with HTML you can just edit the source in a text editor); most PDF readers can’t reflow text (don’t think I’ve ever used a tool—reader or writer—that could); most PDFs can’t be edited freely; and PDFs don’t render well on screen.
You’ve thrown out the HTML ecosystem because some people abuse it, and are choosing to use PDF despite the rampant abuse and many problems of the format, because it’s theoretically possible to work around those problems. (You may quibble with my judgement of “rampant abuse”, but most PDFs I encounter on the web are PDFs for no good reason and perform terribly, loading slowly and being worse for task performance. It is perhaps a slightly different class of abuse, since spying on users via JavaScript isn’t part of it, but it’s related inasmuch as it’s not about meeting the needs of the user.) This does not seem internally consistent.
For what it’s worth though...
That 1MB of PDF still loads nearly instantly on my phone - subjectively no slower than the equivalent much-smaller bare-bones HTML version.
PDF patents are all licensed royalty-free for the normative parts of the PDF 2.0 spec.
PDF tooling sucks, but I’d rather see effort put into improving this situation than into yet another expansion of the web browser.
Doing things like jumping to the end takes perhaps 200–400ms to render that page in the PDF, where HTML would be instant (meaning “less than 16ms”).
No way would I access this on my not-overly-powerful phone: Firefox would download the PDF and try to open it in a local app (EBookDroid) instead, so all up I’d expect that’d make it something like 20–30 seconds to load, instead of 1–2. And the text would be minuscule (or only tiny in landscape mode) instead of sanely sized, further disincentive.
Good to know about the PDF patent situation. Do you happen to know how relevant that actually is to PDF tooling? Is it a new version of the file format, or a specification of the existing? (My knowledge of PDF is limited; I know the general concepts and how it’s put together, but not much of the intricate detail or PDF versions.) That is, does this help for existing documents, or are existing documents still stuck in a patent tangle?
On tooling, I just don’t believe PDF tooling is capable of being excellent; it’s a publishing format—a compilation target more than anything else—and by design not conducive to manipulability, where HTML is an authoring format, so you can work with it. Much of the stuff you can do with HTML tooling is by design fundamentally impossible with PDF. They’re very different types of formats.
I don’t think there’s a conflict between PDF being a publishing format and having great tools for producing that format. The editing can be done to an intermediate application file format (e.g. reStructuredText, or OpenDocument Text, or Microsoft Visio), with the result rendered to PDF.
Oh and that’s unfortunate to hear the mobile Firefox experience is poor. I note there’s a project to add pdf.js support - this is the kind of thing I mean by improving PDF support.
The Tj operators (for printing text) are operating on numbers instead of ascii text making it difficult to read.
I'll see what I can do to duplicate your PDF using some cli tools and a text editor... (hopefully I have time to do it today)
FYI the file has already been linearized through qpdf.
Remember that PDF is mostly a text format that's pretty simple at its heart. You can actually hand-write a PDF, though it's painful: https://brendanzagaeski.appspot.com/0005.html
Take a look at the cli tool QPDF. Once you uncompress each of the dictionaries within the file you can load it up into a text editor and see what is possible. Take a shitty large PDF as an example and then one produced by LaTeX. See the difference.
Free tools can and do produce amazing PDFs that are small and quick to render. Honestly, PanDoc is probably the place to start.
Both ecosystems — HTML and PDF — are under rampant abuse, I agree. But with PDF we could end up with something better — Not because it's most constraining or more freeing but because we would be doing the hard work of layout once, on the server, not a billion times on our battery and cpu constrained portable devices.
My conclusion:
- Use the streaming format with the index headers at the start
- Break the document into pages to aid rendering speed while jumping around but use zero margins on top/bottom to hide the fact that they are separate pages
- Do the layout at publish time for a few screen widths — No need to reflow perfectly for every possible width
We can get started with this and then maybe later tweak the PDF format — A text format with a binary format with a text format within a binary format kinda sucks. Should be completely binary. Maybe use an existing serialization format like protobufs.
Also, while not it doesn't have to be compatible with the current web, there should be mechanisms that translate the new web into (maybe limited) form of the current web, to utilize at least some of the tools and standards.
"Netflix needs a DRM module. Download this official release from <Browser Org>, <Netflix>, <Tom who really likes DRM and encoding alrogithms>".
Or am I just too naive and don't know enough about how browsers work?
Perhaps commercial products can improve over the quality of results, but simple text and keyword search can work this way.
I guess search engines could apply a score to each site so that those who spam, get down-graded, but then how would they do the scoring? By crawling? Then you still need a crawler and might as well build your own index.
What do you think?
HTML is the compilation target, the low-level markup language that Markdown, reStructuredText, &c. can render to.
Compare:
<strong>important</strong>
with: @strong{important}
It looks like some characters saved here. I don't know how much savings would be in the long run.For big sections, one could use @begin{blah} content here @end{blah}.
It looks a bit plesant in my eyes (than html). And only _one_ special character @, which is repeated twice to escape it.
>> All you're doing is fracturing the ecosystem needlessly.
I haven't done anything yet ;)
Gemtext might not be as full featured as some might like, but you can always serve markdown instead.
Table and math bring more problems. Probably SVG or other images can solve them, but it is a pain for authors. In my opinion, no plain-text markup language today has tables with features even remotely close to HTML (at least colspan, rowspan and alignment) while remaining convenient to both write and read. AsciiDoc is fine in terms of features, but not in terms of syntax: