Show HN: This website is valid JSON
webdatarender.com
webdatarender.com
- the content-type of the page is "text/html", so the browser is trying to render html
- there is no special meaning of the #render key to the browser (again the browser doesn't know this is json)
- browsers are very fault tolerant so they'll just skip over your document till they find pieces of html
- the JS made by the author is the part that parses the json as json, and it uses the #render key as its metadata section
You can try opening a file like this to test for yourself to get a sense of just the browser-parsing part, test.html:
{
"test": "hello",
"myHtml": "<html><meta charset=utf-8><h1>Hello World</h1></html>"
}
I do however get this warning in FireFox so perhaps this is pretty fragile:> The character encoding declaration of the HTML document was not found when prescanning the first 1024 bytes of the file. When viewed in a differently-configured browser, this page will reload automatically. The encoding declaration needs to be moved to be within the first 1024 bytes of the file.
I'm guessing differently-configured means not in quirks mode? Does quirks mode just make the browser extra fault tolerant?
> I'm creating a blog platform using this concept.
Yeah, please don’t; this is an abomination that’s fun for demonstrating and teaching how these things work, but should absolutely never be used in reality.
You could use this as your data model foundation upon which to build a generator, but you should never under any circumstances actually serve this stuff.
New ideas seem like bad ideas, because otherwise people would be doing them. One could imagine your objection applying to React: "What a horrible idea. It's a nice hack, but under no circumstances should you actually build webpages like this. Webpages are built out of HTML and valid JavaScript, which this is not..."
The "fun" of this site is that it's its own API. Sure there are better ways to accomplish this but abusing quirks for fun and profit is the hacker spirit.
This is a good hack. It brings a smile to the face and warms the heart (if a good hack is what you were looking for). And I think that was the point anyway, not to come up with a professional architectural pattern.
So long as the JavaScript executes, this doesn’t actually harm accessibility: as it loads, it slurps the JSON, and turns it into a perfectly normal web page; rather like XSLT, as others have pointed out. And quirks mode isn’t that serious a problem. It’s a mild nuisance at most, really.
But I do agree with your last sentence.
The web is pretty much best-in-class for accessibility matters. (There are a few isolated cases where native desktop or mobile apps can do better, mostly to do with efficiency.) HTML elements have defined semantics, so that things like headings and links are automatically navigable, and sections, headers, footers and navigation lists become waypoints. Then ARIA attributes can be used to provide any further metadata necessary, such as to mark up a tabs widget to show how to interact with it. And that’s still key—accessibility needs to care about interactions (which tab is open? and did the content available change?), so state matters. Thus, accessibility tools will never care about any format that you are projecting from, like this JSON; they must only care about what is materialised, which is the HTML DOM. (Besides all that, the only sort of “consistent JSON format” that you could have would be basically an encoding of the HTML, which would be verbose and subjectively ugly compared to the HTML serialisation, e.g. ["a", {"href": "/"}, ["Home"]] or {"tagName": "a", "href": "/", "children": ["Home"]} instead of <a href="/">Home</a>, and miss the whole point here that the JSON is representing data rather than what the user sees.)
If you’re not familiar with accessibility stuff, I heartily recommend looking into it. If you can, find a blind person and see if you can watch them using a computer or phone. It’s really fascinating (I’ve never seen anyone be bored by it) and super useful if you ever contribute to making just about anything on a computer. Even people making documents in a word processor can learn things like “use actual headings rather than just making the text bigger and bold, because the semantics are useful”.
React is a different approach, but it plays by the rules and doesn't serve invalid HTML or Javascript.
It's a fun hack, although not exactly very unique. See also the very old by now "Website in a PNG file" concept, which does the exact same thing: https://gist.github.com/gasman/2560551
React and its ilk were designed for use in apps, where there are meaningful advantages in powering things entirely with client-side scripting rather than generating server-side HTML and possibly enhancing it on the client side with scripting.
They were then abused by increasingly many people for rendering static content, things like blogs.
These people were using the wrong tool for the job, and it has had a detrimental effect on the web.
Now, pages are regularly far heavier than before with expensive client-side code to do stuff that should have been done server-side in almost all cases, and it became popular enough that search engines eventually had to cave and introduce a full JavaScript execution environment in their indexers, tooling everywhere got a lot more complicated, and the last state of the web was worse than the first.
There is a place for things like React: in rich apps, and perhaps even for server-side rendering, though I’m not fond of that for things like blogs because it encourages you to end up depending on it on the client side too.
But client-side JavaScript app frameworks made some formerly-impossible things possible, and formerly-complex-and-unmaintainable things tractable.
This monstrosity, on the other hand, offers no actual benefits for the user, and does introduce a few new problems (quirks mode for styling, and an unnecessary dependency on JavaScript). And so I say it should never be exposed to the client side. Use it as an input format for your blog generator if you like, but don’t try shipping a wonky JSON/HTML polyglot directly.
As for your React commentary, it's 4:42am my friend, and I was just sad to see someone take such a hot steamy dump on someone's work on a Show HN thread without a single other person standing up for them. But all of your points about React can be summed up as "well, yes, that's what happens when something is successful: history is rewritten to make it seem like it had a place from the beginning."
If someone was like "Show HN: React - a new way to write websites," it feels like a guarantee you'd be right there like "But it breaks when you turn off Javascript!" Meanwhile, even Tor admitted defeat long ago and enabled JS by default.
Of React (and its ilk: React was by no means the first project along these lines; as an example, I can think of having hit a couple of full JS-required Knockout sites well before React was a thing, where they would have been better as prerendered HTML), I’m not saying that history was rewritten to say it had a place from the beginning, but rather that there was a place for it from the beginning: that there is a certain type of app where there are very substantial benefits for the user in doing at least some parts on the client side (that was where the jQuery style of progressive enhancement started, and then things like Backbone steadily expanded it), and then that architecturally there are substantial benefits to going all in on client-side rendering if you need this sort of enhancement (this was what Knockout tended towards, and what ExtJS and React more fully realised)—but that this had costs, too, in that it broke the traditional model, making life harder for all kinds of tooling and making pages heavier, so that it shouldn’t be used everywhere.
React was not initially intended as a way to write web sites, but rather web apps. It’s an important distinction. For apps like Facebook and Twitter, the advantages of server-side rendering were not so applicable, and the benefits of full client-side rendering more marked. Unfortunately, the SPA craze grew further, and people liked using one tool everywhere, and so it became more and more common to use React in places where it was inappropriate at the time; until finally tooling like search engines caved on the whole JavaScript thing.
I still don’t like how often I find normal websites depending on JavaScript for fundamental rendering (it may not surprise you to discover that I browse with JavaScript disabled by default—mostly for performance and minimisation of annoyances), but at least React has purpose and some advantages, and has done from the start.
Meanwhile, this thing here doesn’t get you any of the benefits of client-side rendering (things like lighter and faster subsequent page loads and transitions, by using AJAX and semantic knowledge), but does carry all of the costs of client-side rendering: poorer performance, and making life harder for tooling of all kinds.
It's a good thing people exist that aren't deterred by gatekeeping comments like these and actually try to innovate or simply play around and have fun with programming.
If you can retain full functionality (and hack on your ideas) AND be standards compliant (to make sure someone who decides to start offering a new web browser doesn't have to worry about 10% of websites serving this instead of valid HTML), then you should do that.
I think the thing that's annoying is trying to sell other people on using such a tool. Anyone crazy enough to try it should jump right in, but don't try to talk people into it as an actual good blog option.
I call this sort of project "linux on a wristwatch." It's totally cool and a fun hack, but there's no real utility to it beyond an art piece.
Or actually, you can probably just put some closing angles before the start of your render to make sure you don’t get broken from above. Little sketchy though.
The script might be able to set up mutation observers that notice stuff being added to the DOM and immediately remove it into a buffer the script maintains. This might actually be pretty viable.
> Or actually, you can probably just put some closing angles before the start of your render to make sure you don’t get broken from above.
That won't help with the fact that the data will actually be "corrupted" by the HTML parser. This is why all the HTML inside the JSON in the example is HTML-encoded.
One plausible fix for _that_ is to have a <plaintext> tag right after your <script>. So put the #render as the first thing in the JSON, set up mutation observers in the script, <plaintext> to prevent HTML-parsing of the rest of the doc, and this might be pretty robust to random HTML bits in the JSON data.
That seems more plausible since `<plaintext>` can't be closed. In fact mutation observers can be used to get and process the partially downloaded JSON before onload (I'm pretty sure this is possible but haven't tested, YMMV). That would be still horrible as a general solution, but might be actually an interesting solution for more limited situations.
It may not significantly affect users of modern browsers in their default configuration with good internet connections, but it does affect plenty of other things.
You want to parse things in the document? Now you need a whole different suite of tools from the usual tools you use. Your library that parses all the meta tags, identifies the content, &c. is now useless. Now you need either a full JavaScript execution environment, or a JSON parser instead of an HTML parser (and that JSON has completely lost the semantics that HTML provides, so you can’t query things like “document title” or “meta description”).
You have JavaScript disabled? Here, have a mess that, well, it’s better than most client-side rendering things in that the content is still probably there, rather than the screen just being blank, but there are reasons why you should always prefer server-side rendering for things like blogs.
You have a slow or unreliable internet connection? Now the page is taking longer to load, and until the JavaScript loads, the page is empty—and it may fail to load.
I object to people doing things like this as more than a fun technical demonstration because it does harm some users.
They're no less useless than for SPA's that render their content in JS. That's all this site is. If your web scraping suite can't handle content loaded from JS then you're already locked out of most of the web.
Same if you disable JS. Most of the web will be broken for you and it takes an annoying amount of effort enable the minimum necessary scripts.
The key thing here is that SPAs normally get some kind of interactivity benefits from being written in that style (though I confess they break things that the platform provides, by reimplementing them badly, at least as often), such as loading same-site links faster. But this thing doesn’t do that; it’s purely a projection, like XSLT. It should be done as part of a generator or server, rather than on the client side.
I can imagine the author had their fun, I also had fun reading the article, but I don't think this is leading to an accessible internet.
Interesting idea, but Google has a dominant position in both the browser space and the search-engine space. They can punish such pages in the Google search rankings, and this avoids the obvious retort of Why are you deliberately making my browser worse?
Would you say the same about taking a browser -- which was designed to be a document viewer for researchers -- and turning it into an entire application execution platform like we have now?
Point is: It's silly to say things like this, because this is how innovation happens.
Look at how many over-engineered platforms/frameworks Microsoft made, and it just ended up being over-complicated or reached a point where it was no longer worth continuing.
The only tenuous claim it can have is that the document is JSON, so if you want to parse the data, maybe you’ll find it easier? But in practice this is not useful: you can already embed or link to a JSON representation in the HTML, and that JSON representation then won’t be constrained by having to embed the renderer, either.
Not all inventions are useful.
It is not even vid JSON response to begin with (it's served with text/html), and if the fetcher is configured to ignore that, you might as well configure the origin server to serve the JSON response based on the accept header.
# Markdown header
## Subheader
### Section header
1. Numbered
1. List
- Unordered
- List
[//]: # (<html><body></body><script src="https://cdn.jsdelivr.net/npm/marked/marked.min.js"></script><script>var doc = document.children[0].textContent.split('\n'); md = doc.slice(0, doc.length - 1).join("\n"); document.body.innerHTML = marked(md);</script></html><!--)
That last line has varying degrees of invisibility in different markdown viewers I looked at.
Obviously there's optimization that could be had here but this is equally hacky IMO and simpler since you can just write MD instead of JSON. You lose a couple of key features though -- templated components for example. However, I imagine you could shoehorn those in without much effort.This has the advantage that without JS enabled you'll get poorly formatted markdown in the browser.
data:text/html;charset=utf-8,%7B%20%22foo%22%3A%20%22bar%22%2C%20%22baz%22%3A%20%5B%20%22qux%22%20%5D%2C%20%22%23render%22%3A%20%22%3Chtml%3E%3Cbody%3E%3Cdiv%20id%3Dcontent%3E%3Ch1%3EMy%20fancy%20document%3C%2Fh1%3E%3Cp%3EThis%20is%20a%20completely%20%26quot%3Bnormal%26quot%3B%20HTML%20document.%3C%2Fp%3E%3C%2Fdiv%3E%3Cstyle%3Ebody%20%7B%20visibility%3A%20hidden%3B%20%7D%20%23content%20%7B%20visibility%3A%20initial%3B%20position%3A%20absolute%3B%20top%3A%200%3B%20%7D%3C%2Fstyle%3E%3C%2Fbody%3E%3C%2Fhtml%3E%22%20%7D
You can paste this directly into your URL bar, and view source to see the valid JSON document.What the submitted link achieves is impossible without JS, which is also why it's a cute hack and should only be used in production with the knowledge that it's not going to work well for clients without JS.
data:text/html;charset=utf-8,<meta http-equiv="refresh" content="0;url=http://google.com/" />
An oft-unheard voice of wisdom and sane engineering practices speaks. Hark, ye mortals.
I have JS disabled by default. Until I read your comment I was just confused what was going on.
data:text/html,{"data":"Hello.","whatever":"<script>onload=function(){document.body.innerHTML='<h1>'+JSON.parse(document.body.innerHTML).data+'</h1>'}</script>"}I put the meta tag to render some special characters without relying on the server, but HTTP header would be a better option (Content-Type: text/html; charset=utf-8).
And you nailed the process.
Also, if you are interested, the project is on GitHub: https://github.com/webdatarender
In particular browsers tend to be a lot more enthusiastic about treating local files as unicode than they are for responses from web servers.
The reason the meta charset doesn't pick up is you're missing quotes around `utf-8` (also it may be past the 1024 byte mark, hard to say without running wc).
For the third one, there's no "skipping over" involved. The <html> and <body> opening tags are optional in HTML. The browser sees some non-whitespace text (the opening '{') not in the context of any other tag, auto-opens those tags, and the text starts being parsed as body text.
If the page loaded slowly enough, over multiple packets that arrived separated in time, you would see the JSON text coming in and rendering as text, with all the curlies and all, until the browser gets to the <html hidden> part of the JSON. At that point, some common-error fixup dating back to the Netscape/IE 3 days kicks in: when you see an <html> or <body> tag and one has already been opened (due to some stray content at the beginning of the file, likely), you don't open a new one, but copy the attributes to the existing one. This copies the "hidden" attribute to the <html>, which hides it. After that either the script executes and does its work, per your fourth bullet point, or script is disabled and the style rule inside <noscript> unhides the <html> element, but at that point you will of course just see the JSON text, parsed as HTML.
> I do however get this warning in FireFox so perhaps this is pretty fragile:
I'm not sure whether the site changed since then, but now it's sending `charset=utf-8` in the content-type header, so there's no meta prescan at all. Which is what allows the emoji after "Need Help?" to be decoded correctly.
Were you seeing the meta prescan warning on your local test file, or on the site itself?
> I'm guessing differently-configured means not in quirks mode?
No, it means things like preferences about how to handle pages without encoding declarations via various heuristics (e.g. scanning byte value frequencies and guessing what character encoding is in use based on that).
Quirks mode is determined solely by the doctype of the page. This page has no doctype (since it starts with a '{'), so it's always in quirks mode. And what quirks mode does is enable a set of behaviors designed to not break content written more or less in the "before HTML 4.0" era. https://quirks.spec.whatwg.org/ should have a more or less exhaustive list of quirks mode behaviors browsers are expected to implement. Some of these are extra-fault-tolerance (e.g. the hashless hex color and unitless length quirks), while some are just replicating pre-CSS browser rendering behavior (like the line height calculation quirk, which a bunch of sliced-image stuff relies/relied on).
> The JSON must contain only pure information without any concern about design or markup.
Wellll.. I mean the json has markdown syntax in it. That's a lot nicer than html tags for markup, but it's still markup.
In case anyone reading hasn't seen it before, browsers have a thing called XSLT built into them that does something similar for XML documents. You serve up your data as XML, add a tag that points to an XSLT document, and the browser will use the transform on your document automatically. If the resulting output is renderable XHTML, it'll display it like a regular webpage.
That being said, I wouldn't recommend doing that, since XSLT is a programming language whose syntax is XML itself. Apparently web standards folks in the early 2000s thought the future was XML all the way down.
It's very cool, but also, very strange, if you come from a procedural language background.
Also, XSLT is really given power by being mixed with XPath (which is procedural).
I suspect that if I had known more about FP back when I learned it, I would have had an easier time. I wrote up a really long, painful post about XSLT, way back in the early 'oughts.
XPath is not procedural. It is entirely composed of expressions that compose and return results. The closest common language might be SQL.
It was pain and suffering developing these style sheets, but once done it worked like a charm.
Was Adobe FrameMaker involved?
2.0 was (maybe still is?) supported by Saxon, a proprietary library, needing to be licensed. Built-in support was capped at 1.5.
If I remember, in order to run XSLT 2.0, you had to license and install Saxon on your server (i believe it was a Java application, meaning the server need to have a JVM), then call it, using system calls. Awkward, at best.
But this was many moons ago. I suspect/hope that things have improved, since then.
XQuery 3.1 has multiple open source implementations: Saxon, BaseX, xidel, and xqerl.
There is a pure functional programming language waiting to be resurrected from the ashes of XSLT, following Haskell's original idea that a program is a pure function from a stream of inputs to a stream of commands. I just hope that someone removes the XML from it, to save on the pain of my pinkies mashing those greater-than and less-than signs. I like the abstract idea that XML is recursively embeddable, and even the radical suggestion that maybe XSLT should be involved in metaprogramming its own syntax trees, but I am not sure that this is worth it. Maybe if we blur the lines between our code editor and a generic UI, someday?
Even Go, which is really not amenable to any sort of in-language DSL or syntactic sugar, is nicer to use than XSLT itself.
XSLT locks you in a trunk without enough tools to do what you need. I remember using a variant of XSLT Microsoft put out back in the day that let you embed Javascript into it... but then I noticed, why not just do it in Javascript? So I did. And it was soooo much better.
There are good ideas in XSLT, but they are so buried in a cascade of bad decisions and limitations and restictions and missing functionality and bizarre ways of doing things (and I mean, I speak Haskell, pure FP doesn't scare me, and they're still bizarre) that the best remedy is just to start over.
That's XQuery. Version 3.1 handles JSON elegantly.
At NLnet, we use it as a static site generator.
Apparently there's an online playground [1] and Beta support for LSP now [2]. I've only ever used it in a containerized-Java setting.
[0]: https://mjdresdner.medium.com/dataweave-2-x-playground-comin...
[1]: http://dwlang.fun/
[2]: https://github.com/mulesoft-labs/data-weave-language-server
Pure functional programming language, where every expression is a generator. And it's installed on a ton of boxes already.
For example, if you build a social network profile page using XSLT, alice.xml and bob.xml can reference the same profile.xsl stylesheet which converts the XML profile data to the rendered HTML page. Since the profile.xsl stylesheet can be cached, whenever a new profile is visited, only the profile data in XML needs to be transferred.
Note that the templating can also be used to save data when rendering a single page. Imagine a Twitter clone that displays a list of tweets. If, hypothetically, 80 characters of Tweet data plus 500 characters of XSL template, produce 250 characters of HTML markup, you only need 3 tweets to start saving data (250x3=750 > 500+3x80=740).
It should have been, but just like the masses rejected LISP for its parens, the masses rejected XML for the closing tag. We could have avoided PHP and the zoo of MVC frameworks.
External entities, namespaces within namespaces, CDATA, namespaces that confusingly look exactly like URLs, but aren't, parser vulnerabilities.
The bad smell was mostly coming from all the things inherited from SGML - CDATA, entities etc. Although I would also add that separation into elements and attributes is not a particularly convenient way to model many things; it mixes up differences that should really be orthogonal (ordered/unordered, atomic/composite etc).
In my schema, I have a <foreignDocument> element, whose contents is defined to be a foreign document.
Yes XML is more expensive to parse, but so is all the extra stuff you have to do to get around the various limitations of JSON, which could have been done cheaply by an optimized XML parser.
I doubt that. Javascript frameworks have their value inside a corporation - they enable replaceable cog worker programming. The result sucks, but hiring and replacing React programmers is easier.
I don't think they actually do anymore, not all of them for sure.
> web standards folks in the early 2000s thought the future was XML all the way down
Yes. What fools, to think that machine-produced content transmitted from machine to machine and displayed throw other machines would be better handled and more accessible if formatted in machine-friendly ways that were still readable to humans...
But it wasn't to be, not just because XSLT was somewhat botched (way too hard, way too many enterprise-looking warts...) and probably insecure, but because people are too lazy to produce readable markup (who wouldn't get tired of reading <prst><artdd><!CDATA[[ all day long...) and to make sure tags appear only where they should and get closed properly. Add the inevitable touch-of-death of enterprise vendors ("let's use xml in such a way you'll only be able to deal with it with our tools!"), and the game was up.
Isn't markdown literally just HTML shortcuts for a limited set of common markup cases? IIRC, it's not uncommon to have to insert HTML tagging into markdown to get certain things to work.
Assuming it's a superset though, the part I was talking about above was the (markdown - html) part, where markdown adds less noisy syntax for bolding, headings, lists etc.
I think it was called "Where," or "There," or something.
The only advantage I see is that you can actually support markdown like you do in your page, since it's less verbose than HTML-tags and doesn't make the text unreadable if you read it as plain text.
But with HTML5 you can introduce arbitrary tag names to pretty much get the same with an XML-like structure instead of json. Just use <subtitle> <what> <description> and supply CSS for it. If you want to consume it with something else, swap out the JSON parser for an XML one, navigate to the body tag and from there on it's the same thing.
I mean it's a cool trick you came up with, but it doesn't seem worth the effort, and relying on quirks mode seems brittle.
<form method="POST" enctype="text/plain" action="http://example.com">
<input name='{"key1":"val1","params":{"input":"value","list":[],},"dummy":"' value='"}' hidden>
<button>Submit</button>
</form>
The important bit is to include a "dummy" key at the end of the JSON object, and an input value that closes the quotes and any open curl braces. That way the "=" character sent in the encoding of the form elements doesn't interfere with the meaningful JSON content.There might be a clever way to get it to submit dynamic JSON that changes based on user input without JavaScript, but I haven't thought enough about it.
This technique is sometimes useful for CSRF attacks.
{
"key1": "val1",
"params": {
"input": "value",
"list": []
},
"dummy": "="
}
If the JSON is just a hidden form value as you suggest, the request as a whole will not be treated as JSON data. Then invalid characters will (usually) be added to the request body by the browser, and the server will (probably) be unable to parse it, causing the request to fail. This is due to how forms are encoded for POST requests.On the other hand, if you're wondering why anyone would ever do this, then I do not have a good answer for you :)
{
"content": [
["html", {},
[["head", {}, []],
["body", {},
[["h1", { id: "main-header" }, [
"Welcome to my",
["span", { class: "red-text" },
["PAGE"]]]],
["h2", { id: "sub-header" }, [
"It's so cool.",
["br", {}, nil],
"Don't you think?"]]]]]]]};Example of a schema: https://github.com/ProseMirror/prosemirror-schema-basic/blob...
Example of a document converted to JSON: https://tiptap.dev/export
[0] escherize.com/w/hiccup.space
[1] escherize.com/w/cljsfiddle
I didn't create the project mind you, it was some NPM package back when I was a fresh 'un who didn't have thoughts on using a billion subdependencies
Anyway, I was nearly broke and when I went to apply for the benefit, I was asked for my resume as a word doc
Clearly, you can imagine how this conversation went down but I had brought a PDF on a USB which I offered to print out instead.
The clerk refused to let me plug in my USB for fear I was going to "hack" her and the HTML page, saving as a PDF with Chrome, took some convincing to ask her to navigate to so it could be printed
I guess the lesson here is that if you expect to run out of funds, make sure you have your most essential documents stored via Microsoft Word?
I have my resume defined using this schema, which is handy because it's supported enough that every few years when I actually need to update it, I can find a website (such as https://resumake.io) where I can just copy in the JSON and get a nicely formatted PDF out of it.
Microsoft Word is fine too but I would still use something like PDF for export and sharing (still not sure how Microsoft Word solves the usb problem for you).
Generally though I like something text-like that I can revision control more easily/edit in vim. Maybe I’ll switch to something like pan doc going forward (then you can generate word if they really want it).
If you had this experience again, there would likely be some other barrier.
Most of the big MVC frameworks offered this out of the box at some point, which made life easier before dedicated RESTful APIs became a thing.
This is a nice project, but it isn't valid HTML.
https://validator.w3.org/nu/?doc=https%3A%2F%2Fwebdatarender...
Any feedback is appreciated!
An example of the output is at http://trout.me.uk/perl/plx.html (that one's actually served over http, it's a demo, please don't blindly run it).
Feel free to drop by #perl and/or #mojo on freenode if you're interested in chatting more :D
If you were writing a desktop application, you could be saving user data as this JSON-HTML quine (JHQ?), so it's automatically accessible on any device, but also still accessible to e.g. the jq utility.
Since any modification is liable to move the "_" field, you might address that by using an array envelope instead; [{"real": "data"}, "<html><footer>"]
For a production site, I think you ought to, server-side, check the Accept header and pre-render the conversion.
Another note: if I do Save Page As in Firefox, it saves the page as standard HTML, losing the JSON.
The page is rendered in QUIRKSMODE because the source is missing the <!DOCTYPE html> at the beginning of the document, to resemble valid JSON.
Well, at least that's clever!
JSON.parse(document.body.innerText);
document.body.innerText=""
etc...Also, send me a link, I would love to check it out!
> The JSON must contain only pure information without any concern about design or markup. All the design and markup required to render the page must be inside the rendering script.
The source code has markup in the form of Markdown, e.g. ## or * text
For example, the basic render "# Title" creates an h1 header, but if you created the render, you know the internal schema of the JSON, so you can use the correct HTML element for that without recurring to this hack.
But the principle is important: try not to use markup on the JSON, because this would make it difficult for others to consume your website with JSON.
Since there is no standard for representing rich text in JSON, if you want rich text, this is simply an unavoidable problem. At least Markdown is semi-standard and people can get libraries for it. Embedded HTML would also work. Defining an ad-hoc rich text embedded would be worse.
The index page at this point is mostly just a DOM skeleton on which to hang references to CSS, media, scripts and metadata. We might as well cut out the last step already.
If there was an official HTML-to-JSON format, we could even use that as well. It would probably already exist, but the question is always what to do with node attributes, text nodes and child-nodes. There's a dozen ways to organize them in JSON.
tldr: the html is stuffed in the exif data!
[1] https://gist.github.com/gasman/2560551
[2] https://github.com/codegolf/zpng
[3] https://xem.github.io/terser-online/ (If you pick the packing method "Zopfli (DEFLATE)", open the zopfli options and change the format to "zpng", you can write directly to the "Minified (Terser)" input to download the optimized PNG. Yes I wrote that part of code and that required way too much algorithm: https://github.com/xem/terser-online/blob/5cc33125/compress....)
Now that I have built and have played around a bit, I don't see how can I build a website in any other way.
For me, the website's information is important, not the design. WDR exposes that.
Did you play around with just using markdown directly and embedding html tags in that in a similar way to format the page? I can't imagine myself intentionally writing web content using JSON. I'd probably have some other format up front that would be converted to JSON, but at that point I may as well just write some static page generator to create the html.
I think a hard problem ahead would be how to convince other people to also adopt this framework, and how to make sure that people are using the same JSON keys to mean the same thing. I'm a little skeptical that this can happen on its own, because we've already tried a standard that was supposed to designed to convey just the content of the website without the presentation -- HTML.
I'm curious how you plan to prevent your JSON format from being (ab)used for presentation purposes, where people add extra content to make the page display a certain way. And if you have a plan, is this plan feasible using HTML as well?
At first, I was thinking to use some kind of general schema, but those things never work. So I decided something much simpler: the render determines the schema. This is an important aspect that I am working on. If you choose some specific render, your JSON will have a specific shape. For example, if the render is for a landing page, the JSON could be something like "about", "products", "team", etc. (never <section>, <h1>, etc...). But it is too soon to tell I'm still thinking about it.
(but an interesting corollary is that the render could be also a program with the "config" part. It would be an alternative to website builders. Every render would be a different WB tailored for the website)
> I'm curious how you plan to prevent your JSON format from being (ab)used for presentation purposes, where people add extra content to make the page display a certain way. And if you have a plan, is this plan feasible using HTML as well?
I don't have a plan for that yet, and it will be difficult because we are accustomed to mixing data+markup. But I think the mindset is to stop thinking the website for just the browser. The website could be information first, presentation later.
However, you need to make sure the platform can eventually catch up with the de-facto standards people are coming up with, so that people don't keep having to keep on re-inventing these wheels. As a comparison, see how JavaScript features evolved from jQuery and Underscore features.
> But I think the mindset is to stop thinking the website for just the browser.
I agree. I'd suggest that you think through how you can convince people to stop thinking about the website as just for the browser.
In the last decade, many people might have thought that the rise of the mobile web would made web developers separate presentation from content. But instead of using the same HTML with different stylesheets, what actually happened is every major website started maintaining two completely separate websites, one for web and one for mobile! I'd recommend thinking about why this happened, to make sure your framework doesn't nudge people down the same route.
Again, wishing you success!
* We have an XSL stylesheet that ignores input and renders out the HTML contained in "#render"
* Along with the JSON file, we send a header `Link </stylesheet.xsl>; rel="stylesheet"; type="text/xsl"`
* We send the JSON with `Content-Type application/xml`
* Browser renders the HTML via the stylesheet and then the JavaScript takes over and renders the page without quirks mode
Sadly, this doesn't work because when the browser can't parse the JSON as XML, it stops processing and doesn't call the stylesheet :(
HTML1527: DOCTYPE expected. Consider adding a valid HTML5 doctype: "<!DOCTYPE html>". webdatarender.com (2,0)
HTML1513: Extra "<html>" tag found. Only one "<html>" tag should exist per document. webdatarender.com (86,15)
SCRIPT1002: Syntax error render-basic-1.0.3.js (1,4)
HTML1506: Unexpected token. webdatarender.com (86,104)
I tried to solve the quirks mode, but it is kind of hilarious that the browser needs the odd <!doctype html> to flip to standard mode. There are ways to inform the doctype on the headers or XHTML, but this would complicate a simple solution. But I was quite surprised how consistent is the look on different browsers, at least in the modern ones.
I'll accept that answer. As controversial as this comment may be, XML/XSLT is a very good fit for this purpose. It might not be particularly modern, however it's almost universally supported.
Cool, but why?
> An object is an unordered set of name/value pairs.
But this site seems to be assuming that the key/value pairs are parsed in their original ordering for everything to display properly. Is it safe to assume JavaScript will always parse the keys in order?
And then JSON.parse is defined as doing the obvious and only sane thing, setting each property as it goes.
Consequently, so long as your keys are not numeric, yes, JavaScript guarantees that it will all be in order.
But that’s for JavaScript. It is incorrect to treat JSON objects as ordered, because various libraries in various languages will discard the order for various reasons (e.g. efficiency, DoS resistance), since it is defined as being not significant. If you care about the order of things, use an array instead.
Technically, no. Functionally, probably yes.
Any browser vendor that decides to muck with the current status quo would likely break enough code to be a non-starter.
You could use this to syndicate blog posts, sort of like RSS, except each entry is JSON and could be viewed as it's own page rendered by it's own renderer (carried by a CDN and cached by the browser), or by the renderer of the users choice.
I know it's a complex idea, and it's hard to grasp ideas when they're first proposed, but the payoff is often worth the time invested. Even though it's only been <checks notes> 24 years since the RFC defining the Accept header was published, maybe it'd be worth spending say, 10 minutes reading about what it is, rather than however long it took you to write this abomination?
My immediate thought for a use case is to debug APIs. Pass in a param that adds this to the response and get it in a more human readable form. Will test it out.
Wasn't XHTML (and thus html5) supposed to be parsable, since it's basically XML (specifically, the html spec redefined as XML) ?
https://wgx.github.io/anypage/
(It's the world's worst CMS)
As a blogging platform, please don't do this. It breaks all SEO, microformats, RSS feed discovery, rel=me ties, and much, much more.
HTML is great. Please use it for your websites.
source: runs a open-web friendly microblogging platform
Nice project though!
Only that my goal was to create a browser for that an alternative to the web.
It would work on the same principle of separating data from information
Meaning: If the javascript (and other content) loaded via #render is highly cacheable, can this lead to pages that display as soon as the JSON is loaded?
That's what HTML and CSS is, no?
```html
{
"#info": {
"title": "WDR"
},
"subtitle": "Web Data Render",
"title": "# WDR",
"what": {
"title" : "## What is WDR?",
"description": [
"This website is a valid **[JSON](//www.json.org/)**!",
"Check the source code. Instead of the habitual HTML and CSS, you will see just a plain JSON with the website's information.",
"WDR is a format to separate the website's **information** and **design**.",
"The website is readily available to be consumed outside the browser via JSON, but also still presentable to users accessing through the web browser."
]
},
"subscribe": "I'm creating a **blog platform** using this concept. Follow me on [Twitter](//twitter.com/gpiresnt) to be notified when is ready! ",
"how": {
"title": "## How it works",
"description": [
"It works by embedding a small initiator at the end of the JSON file.",
"For example, this is a valid JSON and web page:",
"```{\n \"title\": \"Example Page\",\n \"description\": \"This is an example.\",\n \"#render\": \"<html hidden><script src=/render.js></script></html>\"\n}```",
"The script `render.js` receives the JSON as input and is responsible to render the page."
]
},
"usage": {
"title": "## Usage",
"description": [
"First create an HTML file with the JSON information:",
"```{\n \"title\": \"Example Page\",\n \"description\": \"This is an example.\"\n}```",
"Include the initiator at the bottom:",
"```{\n \"title\": \"Example Page\",\n \"description\": \"This is an example.\",\n \"#render\": \"<html hidden><meta charset=utf-8><script src=/render.js></script></html>\"\n}```",
"The next section explains how to create the `render.js`."
]
},
"only-data": {
"title": "** Pure information **",
"description": [
"The JSON must contain only pure information without any concern about design or markup. All the design and markup required to render the page must be inside the rendering script."
]
},
"create-render": {
"title": "## Creating a render",
"description": [
"Create a new javascript project and install the package `wdr-loader`:",
"```npm install wdr-loader```",
"Call the `loader` function to retrieve the JSON:",
"```import loader from 'wdr-loader';\nloader(data => render(data));```",
"Create a `render` function to handle the data and render the HTML. Below is an example using simple `innerHTML`:",
"```function render(data) {\n document.head.innerHTML = `\n <meta charset=\"utf-8\">\n <meta name=\"viewport\" content=\"width=device-width, initial-scale=1\" />\n <title>${data.title}</title>`;\n\n document.body.innerHTML = `\n <section>\n <h1>${data.title}</h1>\n <p>${data.description}<p>\n </section>`;\n}```",
"The `wdr-loader` code is available on [GitHub](//github.com/webdatarender/wdr-loader) with an example."
]
},
"basic-render": {
"title": "## Basic render",
"description": [
"If you don't want to create a render right now, it is available a basic render (used on this very website) to immediate use:"
],
"download": "[render-basic-1.0.3.js](//webdatarender.com/dist/render-basic-1.0.3.js)",
"instructions": [
"Just download the script and include it on the initiator directly:",
"```{\n \"title\": \"# Example Page\",\n \"description\": \"This is an example.\",\n \"#render\": \"<html hidden><meta charset=utf-8><script src=/render-basic-1.0.3.js></script></html>\"\n}```",
"The code is available on [GitHub](//github.com/webdatarender/wdr-render-basic)."
]
},
"remarks": {
"title": "## Remarks",
"description": [
"• The page is rendered in [quirks mode](//developer.mozilla.org/en-US/docs/Web/HTML/Quirks_Mode_and_Standards_Mode) and can present some layout differences on different browsers.",
"• Although javascript is necessary to render the page, most search engines, like [Google](https://developers.google.com/speed/pagespeed/insights/?url=webdatarender.com) or [Bing](https://www.bing.com/webmaster/tools/mobile-friendliness), will be able to read the page correctly.",
"• If you want to display the JSON for users that have javascript disabled, you can include `noscript` at the initiator: ```<noscript><style>html{display:block !important; white-space:pre}</style></noscript>```"
]
},
"support": {
"title": "## Need Help? ",
"description": "If you need help, have any feedback or just want to say hi, send me an [email](mailto:gpiresnt@gmail.com)."
},
"about": {
"creator": "Created by [@gpiresnt](//twitter.com/gpiresnt)",
"logo": ""
},
"#render": {
"_": "<html hidden><meta charset=utf-8><script src=/dist/render-basic-1.0.3.js></script></html><noscript><style>html{display:block !important; white-space:pre}</style></noscript>",
"css": "css/main.css"
}
}
```But it would be even better if they could get the information directly on JSON. So much carbon saved ;)