SVG: The Good, the Bad and the Ugly
eisfunke.com
eisfunke.com
What's even worse is that because XML is a garbage fire of a format, SVG actually builds its own micro-DSL to work around limitations.
So for instance you have something like:
<circle cx="10" cy="10" r="2" fill="red"/>
Ok, fair enough. Readable, self-describing.But then you have:
<path d="M 10 10 H 90 V 90 H 10 L 10 10"/>
Oh. You can almost hear the designer saying "well fuck it, we'll just jam it all into a string.The fact that SGML-based formats managed to become as popular as they are is really proof that there's something fundamentally rotten in software engineering.
I do quite like SVG though. I mean, if the alternative is drawing with CSS, it's pretty amazing. I don't really agree that it's hard to generate programmatically either, you just have to decide what subset of the format you need, look at what inkscape generates and go from there. If you need to generate basic shapes it's fairly straightforward.
HTML is a pretty good syntax though. SGML-like syntax is just less appropriate for formats which are not text-centered, since the element/attribute distinction becomes superfluous noise in the syntax.
But HTML is a markup language, the problem is that for some insane reason a big chunk of the industry at some point decided that XML was a reasonable serialization format. That's where it went really wrong IMO.
<p>First paragraph<p>Second paragraph</p>
P-element can't be nested, so a <p> will automatically close the previous p-element.But this also means it is necessary to indicate the tag name in the closing tag. For example this would be ambiguous:
<div><p>Hello</>
It cannot be inferred if it is the div or p which is closed.XML is of course different. Since all element have to be explicitly closed, repeating the tag-name in the close tag is in principle superfluous. I think it makes it easier to read though, at the cost of more effort to write.
The fact that they are almost human-editable and can be inlined means we are bound to be disappointed when working with them in this manner
Fir example I recently made this 'hack' for gradient transform. even 'hacking' that was easier than maintaining and understanding complex canvas JS - at least to me.
[0] www.holysnacks.us
I do love the mix of css vars into SVG. it makes a lot of sense to me too.
Here's an example: https://svkt.org/~simias/up/20210212-185105_svg.png
All the elements and paths are SVG generated from some javascript. You can interact with the various elements, move them around and the paths are updated in real time.
Of course you could do the same thing with a raster format by using a canvas and drawing a bitmap for instance, but SVG works at a higher level and lets you do these things a lot faster and probably with better performance. You can use CSS-styling, you can register javascript event callbacks etc...
Or to put it another way: if SVG was meant as an opaque machine-to-machine format, why even bother with the insane overhead of XML instead of using some dense binary format?
"The properties of the lines (color, opacity, fill...) can be used to programmatically add functionality..."
Then why bother with the XML? And why use the embedded DSLs in strings instead of the more verbose XML-way.
However, for when you want that there is SNG¹. It can be a surprisingly fun way to perform minor corrections or filters to an image.
{"path": {"d": ["M", 10, 10, "H", 90, "V", 90, "H", 10, "L", 10, 10]}} (path (d (M 10 10) (H 90) (V 90) (H 10) (L 10 10))) (path d M 10 10 H 90 V 90 H 10 L 10 10)
Shore up some nesting for the commands with pattern matching, returning a path struct: This is the TXR Lisp interactive listener of TXR 251.
Quit with :quit or Ctrl-D on an empty line. Ctrl-X ? for cheatsheet.
Garbage collection is on Tuesdays: bring unwanted pointers to curb by 7:30.
1> (defstruct path () name args)
#<struct-type path>
2> (defun parse-path-args (args)
(match-case args
((M @x @y . @rest) ^((M ,x ,y) ,*(parse-path-args rest)))
((L @x @y . @rest) ^((L ,x ,y) ,*(parse-path-args rest)))
((H @x . @rest) ^((H ,x) ,*(parse-path-args rest)))
((V @x . @rest) ^((V ,x) ,*(parse-path-args rest)))
(())
(@else (error "bad path arguments: ~s" else))))
parse-path-args
3> (defun-match parse-path
(((path @(symbolp @name) . @args))
(new path name name args (parse-path-args args)))
((@else) (error "bad path syntax: ~s" else)))
parse-path
4> (parse-path '(path d M 10 10 H 90 V 90 H 10 L 10 10))
#S(path name d args ((M 10 10) (H 90) (V 90) (H 10) (L 10 10)))
Error cases: 5> (parse-path '(path d M 10 10 H 90 V 90 X 10 L 10 10))
** bad path arguments: (X 10 L 10 10)
** during evaluation of form (error "bad path arguments: ~s"
else)
** ... an expansion of (progn (error "bad path arguments: ~s"
else))
** which is located at expr-2:2
** run with --backtrace to enable backtraces
6> (parse-path '(paath d M 10 10 H 90 V 90 X 10 L 10 10))
** bad path syntax: paath d M 10 10 H 90 V 90 X 10 L 10 10
** during evaluation of form (error `bad path syntax: @else`)
** ... an expansion of (progn (error `bad path syntax: @else`))
** which is located at expr-3:1
** run with --backtrace to enable backtraces (path (M 10 10 H 90 V 90 H 10 L 10 10) :color red)
Though probably I would prefer something more nested if possible, so you have a list of instructions rather than needing to parse them out of a flattened list. And I would likely use an alist rather than a plist for attributes like colour or stroke width. path 10 10 M 90 H 90 V 10 H 10 10 L #000 stroke+(x y) would be fine for expressions but weird for lists of things.
path:
d:
M: 10, 10
H: 90
V: 90
H: 10
L: 10, 10
Just beautiful.XML/SGML are a very effective way of representing tree data (again, non-1D data) in 1D strings. And they're wonderfully extensible, while still keeping a well defined schema. Is there an alternative language that is better? How is it better?
<path d="M 10 10 H 90 V 90 H 10 L 10 10"/>
ctx.moveTo(10, 10);
ctx.lineTo(90, 10);
ctx.lineTo(90, 90);
ctx.lineTo(10, 90);
ctx.lineTo(10, 10);
Not much better. <draw id="my-path">
<move-to x="10" y="10" />
<line-to x="90" y="10" />
<line-to x="90" y="90" />
<line-to x="10" y="90" />
<line-to x="10" y="10" />
</draw>
<path draw="my-path" />
EDIT: Come to think of it, the above is probably a lot easier to animate using CSS or SMIL. Animating the `d` attribute on `<path>`s is theoretically possible but in practice weird at best or browser inconsistent or simply impossible at worst.EDIT 2: In D3 it is not uncommon to draw area graphs with a thick border on top. The only way to do this is with two `<path>`s with almost identical `d` attributes. Allowing drawings to be strung together with multiple idrefs could solve that:
<draw id=line-path><!-- ... --></draw>
<draw id="area-box-close">
<line-to x="100" y="100" />
<line-to x="0" y="100" />
<end />
</draw>
<path draw="line-path" class="line" />
<path draw="line-path area-box-close" class="area" />
Maybe this isn’t such a bad idea after all.I SVG is pretty inelegant, but it's pragmatical.
Also: Reminder SVGs goal is presentation, so paths even for a simple monochrome icon have _hundreds_ of control points. Something more intricate or a larger graphic? Thousands easily. All the examples here are toy examples, that don't scale to what SVG is actually used for.
If I want something free-form text based, I've been more happy with e.g. YAML based formats.
If you decided to go fully XML on "<path d="M 10 10 H 90 V 90 H 10 L 10 10"/>" it would probably look like this:
<path>
<point>
<x>10</x>
<y>10</y>
</point>
<point>
<x>90</x>
<y>10</y>
</point>
<point>
<x>90</x>
<y>90</y>
</point>
<point>
<x>10</x>
<y>90</y>
</point>
<point>
<x>10</x>
<y>10</y>
</point>
</path>
Some would say the fact the designers of XML went for "M 10 10 H 90 V 90 H 10 L 10 10" instead shows they thought XML is too verbose.simias likely agrees that XML is too verbose - and wonders why they used XML at all
And it’s not like XML has no good ideas in it either. You may disagree about the execution, but the idea of seamlessly mixing structured data defined by independent standards in a single document is pretty damn powerful.
Also, your example is rather strawman-ish: there is no point in using separate elements for individual coordinates, and anyone designing the format would know this. A more realistic example would have <path> contain an ordered sequence of <move-to x="..." y="..." /> and <line-to x="..." y="..." /> elements. (In fact, SVG already contains <line /> and <circle /> which use roughly this structure.)
The thing is, just having an XML parser isn't enough to parse SVG. You also have to parse a DSL. So I think the argument is "why not just use only a DSL, that does a good job of describing the data model?"
I imagine they thought that the path DSL is simple enough to parse (and I seem to vaguely recall PostScript has something similar), while the overhead of representing path nodes as XML elements would be too high.
I can't help but think s-expressions would have been a better choice.
You say that - but here's some real world XML from the widely used 'GPX' file format[1]
<trkpt lat="47.644548" lon="-122.326897">
<ele>4.46</ele>
<time>2009-10-17T18:37:26Z</time>
</trkpt>
The truth is I could have put the X coordinate as an attribute and the Y coordinate as a child element and it would still have been a fair representation of real-world XML documents.And GPX is one of the better XML formats! You want to see nightmare XML? Go look at SAML.
[1] https://en.wikipedia.org/wiki/GPS_Exchange_Format#Sample_GPX...
JSON and YAML do not make such a distinction.
(I do concur with others who have noted here that the "elements" versus "attributes" distinction makes XML a poor choice for serialization, but XML as a serialization format isn't really the issue here.)
So using XML to represent data structures instead of marking up text can be pretty awkward, inefficient, and nuanced.
XML Attributes are second class citizens compared to TEXT nodes which can contain CDATA, because attributes undergo "Attribute-Value Normalization" -- having their line breaks, entity references, and white space normalized. Newlines are normalized, leading and trailing white space removed, repeating white space replaced with a single space.
SVG path attributes (as well as simple values like numbers, booleans, enums, etc) are impervious to Attribute Value Normalization corruption, because they don't depend on white space being perfectly preserved (by design, of course), so they are fine to put in attributes.
But if you really care about preserving the exact value of a string, like a password or arbitrary string, you should use <!CDATA[[ ]]> in a text node, not an attribute!
I Wanna Be <![CDATA[ https://donhopkins.medium.com/twenty-twenty-twenty-four-esca... ]]>
https://www.w3.org/TR/xml/#AVNormalize
3.3.3 Attribute-Value Normalization
Before the value of an attribute is passed to the application or checked for validity, the XML processor must normalize the attribute value by applying the algorithm below, or by using some other method such that the value passed to the application is the same as that produced by the algorithm.
All line breaks must have been normalized on input to #xA as described in 2.11 End-of-Line Handling, so the rest of this algorithm operates on text normalized in this way.
Begin with a normalized value consisting of the empty string.
For each character, entity reference, or character reference in the unnormalized attribute value, beginning with the first and continuing to the last, do the following:
For a character reference, append the referenced character to the normalized value.
For an entity reference, recursively apply step 3 of this algorithm to the replacement text of the entity.
For a white space character (#x20, #xD, #xA, #x9), append a space character (#x20) to the normalized value.
For another character, append the character to the normalized value.
If the attribute type is not CDATA, then the XML processor must further process the normalized attribute value by discarding any leading and trailing space (#x20) characters, and by replacing sequences of space (#x20) characters by a single space (#x20) character.
Note that if the unnormalized attribute value contains a character reference to a white space character other than space (#x20), the normalized value contains the referenced character itself (#xD, #xA or #x9). This contrasts with the case where the unnormalized value contains a white space character (not a reference), which is replaced with a space character (#x20) in the normalized value and also contrasts with the case where the unnormalized value contains an entity reference whose replacement text contains a white space character; being recursively processed, the white space character is replaced with a space character (#x20) in the normalized value.
All attributes for which no declaration has been read should be treated by a non-validating processor as if declared CDATA.
It is an error if an attribute value contains a reference to an entity for which no declaration has been read.
You should read some of James Clark's criticisms of the official XML Schema (XSD) standard, which motivated him to develop TREX (Tree Regular Expressions for Xml), which he combined with Makoto Murata's RELAX (REgular LAnguage description for XML) to create Relax/NG.
https://en.wikipedia.org/wiki/James_Clark_(programmer)
https://en.wikipedia.org/wiki/Makoto_Murata#RELAX_and_RELAX_...
>Some people, including Murata and James Clark, had critical attitudes toward XML Schema. XML Schema is a modern XML schema language designed by W3C XML Schema Working Group. W3C intended XML Schema to supersede traditional DTD (Document Type Definition). XML Schema supports so many features that its specification is large and complex. Murata, James Clark and those who criticised XML Schema, pointed out the following:
>It is difficult to implement all features of XML Schema.
>It is difficult for engineers to read and write XML Schema definitions.
>It does not permit nondeterministic content models.
>Murata and collaborators designed another modern schema language, RELAX (Regular Language description for XML), more simple and mathematically consistent. They published RELAX specification in 2000. RELAX was approved as JIS and ISO/IEC standards. At roughly the same time, James Clark also designed another schema language, TREX (Tree Regular Expressions for XML).
>Murata and James Clark designed a new schema language RELAX NG based on TREX and RELAX Core. RELAX NG syntax is the expansion of TREX. RELAX NG was approved by OASIS in December 2001. RELAX NG was also approved as Part 2 of ISO/IEC 19757: Document Schema Definition Languages (DSDL).
https://en.wikipedia.org/wiki/Regular_Language_description_f...
https://en.wikipedia.org/wiki/RELAX_NG
https://en.wikipedia.org/wiki/XML_Schema_(W3C)
Schema Wars: XML Schema vs. RELAX NG (1/2) - exploring XML
https://web.archive.org/web/20180429143242/http://webreferen...
https://web.archive.org/web/20180429145711/http://webreferen...
https://news.ycombinator.com/item?id=22756875
>James Clark used Haskell to design and implement an algorithm for validating Relax NG XML schemas (he co-designed Relax NG, and designed its predecessor TREX), to work the ideas out before re-implementing it in (many many more lines of tedious brittle) Java (JING). Haskel works wonderfully as a design and standard definition language, that way.
https://news.ycombinator.com/item?id=25435678
>James Clark's compact syntax for Relax/NG XML schema validation language is quite tastefully designed, an equivalent but more convenient alternative syntax than XML, for writing tree regular expressions matching XML documents. It's way more beautiful and coherent than the official "XML Schema" standard.
[...]
>There's a wonderful DDJ interview with James Clark called "A Triumph of Simplicity: James Clark on Markup Languages and XML" where he explains how a standard has failed if everyone just uses the reference implementation, because the point of a standard is to be crisp and simple enough that many different implementations can interoperate perfectly.
>A Triumph of Simplicity: James Clark on Markup Languages and XML:
https://web.archive.org/web/20130721072712/https://www.drdob...
"The standard has to be sufficiently simple that it makes sense to have multiple implementations." -James Clark
<path>
<point x=10 y=10/>
<point x=90 y=10/>
<point x=90 y=90/>
<point x=10 y=90/>
<point x=10 y=10/>
</path>This could be a better format, but still a lot heaver in term of bandwidth.
<path>
<move x=10 y=10/>
<line x=90 y=10/>
<curve x=90 y=90/>
<end/>
</path> {
"path": [
[10, 10],
[90, 10],
[90, 90],
[10, 90],
[10, 10]
]
}
No need for the x and y, really. And it's still quite legible on one line. {"path": [[10, 10], [90, 10], [90, 90], [10, 90], [10, 10]]} {
paths: [
{
points: [
[10, 10], // 0
[90, 10], // 1
[90, 90], // 2
[10, 90] // 3
],
lines: [
[0, 1, 2, 3, 0] // connect point 0 to 1 to 2 to (...)
],
curves: {
0: [5, 15, 15, 5] // bezier control points for point 0 (x1, y1, x2, y2)
}
}
]
}
The downside (other than the bloat) would be when writing the code by hand, you'd have to keep track of the indices in the `points` array. [ { line: [10, 10] }, { cubic: [ 5, 15, 15, 5 ] }, ... ]
or even [ ["line", 10, 10], ["cubic", ...], ... ]
- anything as long as there's a way to stream the data (which is not possible with your format as you need to receive the whole object before being able to do anything).No, I didn't. I was working from the parent's code. I don't appreciate your attitude though. There's much nicer ways to correct people online.
> No need for the x and y, really
<path>
<m x=“10” y=“10” />
<h dx=“90” />
<v dy=“90” />
<h dx=“10” />
<l x=“10” y=“10” />
</path>
Perfectly sane.Also just in general, SVG's d-path syntax is a fantastic example of smart meeting of and understanding of requirements and users. If they used something more verbose (like `move-to` instead of `M` or whatever), then the file becomes WAY less human readable (too much noise at the XML-level), but more novice friendly. But... will novices need to edit d-paths manually? Of course not! Will they need to edit the XML structure manually? Way more likely. So you can use a more compact expert syntax for paths, because the people who will want to use it will likely be experts. It also means that SVG files are smaller, require less memory to parse, and more human-readable. Seriously great choice on their part.
In XML ID fields are supposed to be unique to the document. Nobody enforces this. Instead getByID returns the first ID. Which means if you ask a DOM element and the root element for an element by ID, you can get two different answers.
Composing multiple schemas into the same file is a wordy, confusing mess, DTDs are straight up broken, and the only time I ever saw someone generate a Grade A XML Schema was by feeding examples into XMLSpy instead of writing their own.
XML asks a question we already knew the answer to: What if we made everyone into a programming language designer? And the answer was "anarchy" because we know that many people cannot design a syntactically consistent language, and hardly anybody can design a semantically consistent one.
James Strachan, as I recall, retired from Groovy before they ever figured out an unambiguous grammar for it. There comes a point where you realize you've made a mess that you are not qualified to clean up (Kernighan's Law). You can either do the very hard work of maturing into the responsibility, or bow out. Strachan is no different than a dozen people I've worked with and countless people I've heard about second or third hand, who I'm more than a little glad moved on.
The guilt is not theirs alone. Part must fall to this shared delusion that you can do anything with software if you only put your mind to it. We have mathematical proofs that tell us that's not true, and yet we still believe in the power of belief. Unfortunately if we didn't believe it a little, then we'd probably never write anything at all, so I'm not sure there's so much a cure as a condition that has to be managed. A little bit can go a long way, and most of us take it too far, sooner or later.
>XML/SGML are a very effective way of representing tree data
Completely disagree:
- It's not effective density-wise because the format is very verbose
- It's not effective parsing-wise because the format is very complicated.
- It's not effective human-wise because you have meaningless distinctions between attributes and child nodes which makes sense for a markup language but not for a serialization format. It's also very verbose which makes it annoying to read and write.
Imagine that you have an object like:
struct Object {
name: "foo",
}
Should you serialize like: <object name="foo" />
Or: <object> <name>foo</name> </object>
And to be clear, that's a rhetorical question. My point is that this distinction doesn't make a lot of sense for a serialization format and forces the developer to make pointless decisions. Here you might say that an attribute makes more sense, but then if later you realize that you can have several names you're screwed, because you either have to come up with your own custom format to store those in an attribute (coma-separated? But what if there's a coma in the name?) or split them off into child nodes.It makes perfect sense for a markup language though, because there's (usually) a clear distinction between the textual and metatextual content. The stuff between the tags is meant to be displayed to the user, whereas the attributes are meant to be interpreted by the machine.
XML just tempts you to use attributes to represent strings, then kicks you in the ass when you least expect it by corrupting your data unexpectedly and unfairly but with full legal authority of the xml standard itself.
I wrote about it in this other comment:
(Note: XAML does have a nice duality between attributes and nodes (sort of) which is kind of interesting. It's been a while since I've used it though, so can't quite recall the details).
You'd think so. While integrating TeX into my editor[0], using JMathTeX[1] as a starting point, I wanted to convert font glyphs to vector paths. At the time, every Java-based SVG renderer (Batik, JFreeSVG, SVG Salamander) suffered from the exact same performance problem: they all used Java's DecimalFormat class to convert floating point numbers to strings. These strings were then concatenated as paths into the output document. Sometimes the strings were re-copied numerous times because of buffer re-allocations and other reasons. (Part 5 of my series[2] on developing a TeX-oriented text editor covers this in depth.)
To get great performance---in Java, at least---even for basic shapes requires a sufficiently large, reusable buffer that doesn't incur memory re-allocations and an efficient algorithm to convert floats to strings---such as Ryu[3]. See my implementation that converts glyph paths to vector paths for details[4]. With these two changes I was able to improve upon JFreeSVG's implementation four-fold.
[0]: https://github.com/DaveJarvis/keenwrite (my editor)
[1]: https://github.com/DaveJarvis/JMathTeX/ (my fork)
[2]: https://bell-sw.com/announcements/2020/11/02/TeXnical-Writin... (my blog series)
[3]: https://github.com/ulfjack/ryu
[4]: https://github.com/DaveJarvis/JMathTeX/blob/d2d678717505765b...
This is an almost universal disease of people trying to write XML grammars (or database schemas before that). Everyone thinks they can unambiguously decode multiple pieces of data stuffed into the same fixed-width field, and pretty much all of those people are wrong. Over, and over again.
That they never learn is a testament to the developer's ability to make fixing their Kruger-Dunning moments Someone Else's Problem. Get promoted or switch teams before the consequences are unavoidable. You can make an argument that even Tim Bray did this.
There are a lot of reasons XML was never going to succeed, and this is far from the most important, but it's a substantial contributing factor.
Array literals in programming languages are pretty much as compact as you can get (start, end, one character element separator), and the type system or a little code often allows you to fix a 1:1 mistake after the fact. This attribute is now a number or an array of numbers.
In XML if you want a properly formatted list you have to use child elements, which gets expensive really quickly so people balk. Especially if you already failed the attribute/child test by mistaking the value for 1:1 when it turned out to be 1:many.
We also already have HTML class, style, and on* attributes as prior art that we use to convince ourselves that one more will be okay. Basically XML gets it from both ends. XML makes entity relationships and irreversible decision, in an era where UML existed and harped constantly on it.
I don't know if there was any fixing it at that point. What I do know is that a lot of us shot meaningful glances at each other and tried as hard as we could to find something else to do until the dust settled. But you really could not avoid XML in the 00's.
I'll throw in my 2 cents here because the responses aren't addressing why the path descriptors are stylised in this way. Whether SVG's underlying document format was XML, JSON, SGML, CSS or whatever is irrelevant. However else you did it, any alternative would be more verbose because the 'd' tag already offers the most compact way of describing vector paths in text form.
And that's the point - it's small. Even in this compact form, it's not unusual for path strings to reach kilobytes in length when floating point numbers are catered for. Thus we need the path descriptions to be as small as possible for optimal parsing of the file and keeping the file size down.
As a side-note I disagree with the author's conjecture that SVG doesn't succeed as a machine-focused language or a human-focused language. I think it serves both sides of the coin quite well in practice (certainly more-so than HTML) and a pared back version already exists in the form of SVG Tiny.
For the most part, SVG path outlines are almost always defined in software like Illustrator and Inkscape. Bezier curves are beyond the average person's ability to create or modify without visual aids. Simple path constructs aren't all that common in practice and the examples I see posted here aren't representative of real-world use.
SVG does provide transform capabilities to move, rotate and scale paths easily, and those features can be applied within a text editor. For anything else there are plenty of visual tools to aid path modification. Personally I've modified many an SVG file manually, but when it comes to the paths I can't say that seeing them laid bare would be of a help to my workflow.
(path (M 10 10 H 90 V 90 H 10 L 10 10) :color red)
It's legitimately hard to beat sexprs for the combination of compactness, generality, and ease of use. If I were to start from first principles in making a vector graphics format, I'd probably make it a schema over EDN. I bet the extra complexity of having explicit maps and vectors would pay off. (path (M 368.296, 1.514) (c -0.908, 0 -1.818, 0.011 -2.716, 0.021)
(c 0.053, 0, 0.106, 0, 0.159,0) (c 147.196, 0.836, 265.306, 94.672, 267.029, 183.784)
(c 0, 0.581, 0.065, 35.087, 0.065, 35.087) (h 49.415)
(v -33.471) (h 37.094) (v -31.376)
(C 719.342, 77.099, 575.826, 1.514, 368.296, 1.514)
(z))Yes an s-expression is imo normally the best representation (more specifically, an x-expression (see Racket) would be ideal). However, practically, it's not actually that useful. It is very useful as an intermediary step.
For example for a tool where the SVG is parsed to an x-expression and cleaned then compiled back. This is one of the [aborted] things I was working on, as the three best existing tools [in JS, Python and Rust] for this all have issues of one kind or another. Or you have a tool that allows you to write x-expressions then compile. But always with this, if that's the case, can't really directly jnterop with the host language/platform -- as you say, would have to be from first principles. And further to that, it would need HTML from first principles because one of the huge, huge benefits is that SVG drops directly into HTML with virtually the same semantics, same styling, same JS interactivity hooks. And the issue is compounded by SVG not just being XML, that it is also a DSL (NB solution is normally to restrict to a subset of features), so the clean structure breaks down pretty often. Imo SVG has to be the way it is, that the tradeoffs of the bastardised XML it uses have to have been made for it to be actually useful. Those tradeoffs also make it difficult to use, and often very difficult to parse easily. But without changing how HTML works, how web browser rendering works, I don't see how there are any alternatives that are substantially better.
At that point, using an s-expression format (x-expression, EDN, I'm agnostic) should be the first choice. One of the first things that would have to be implemented, of course, is a translator to SVG.
My interest in vector graphics is such that, on the few occasions I need them, I just shrug and deal with the format we have.
The circle has unordered properties.
The path has ordered commands. I suppose those commands could be separated and numbered, but isn't that a lot more verbose and less readable?
Instead of SVG and things like FontAwesome in 99% UI cases we use just these
button {
background-image: url(path:M 10 10 H 90 V 90 H 10 L 10 10);
fill: red;
}
or <img src="path:M 10 10 H 90 V 90 H 10 L 10 10"> def port 0,0 100,100:
L 10,10 90,0 90,70 70,70 70,90 20,90 20,70 10,70 10,10
def bank 0,0 400,100:
port 0,0 100,100
port 100,0 200,100
port 200,0 300,100
port 300,0 400,100
def router 0,0 1100,200:
port 50,50 150,150
bank 200,50 600,150
bank 650,50 1050,150
router 0,0 1100,200
I was thinking it would be cool to have an "about" page in openwrt showing your router with all the ports labeled in svg. Should be simple, but I couldn't figure out luci.https://en.wikipedia.org/wiki/Scalable_Vector_Graphics#Devel...
Pragmatically it was a brilliant practical compromise in response to one of XML's horrible flaws as a serialization format (it's way too fluffy), saving enormous amounts of memory and fragmentation by using one string per path, instead of creating thousands of DOM objects with tens of thousands of attributes and text nodes. But you're right, that's the "original sin" of XML that SVG is doomed to suffer with.
PDF is basically just the PostScript imaging model without the Turing-complete stack based programming language.
SVG is just the latest expression of PostScript graphics model (if you ignore all the stuff about fonts), which is based on the Stencil/Paint imaging model with Porter Duff Compositing, and descendent from Interpress and Jam).
The canvas 2d api is pretty much just immediate mode SVG / PostScript graphics, with a few new features added.
JSON's a much better format for structured graphics in some ways, but has its own set of flaws and limitation, like not having comments, and being terrible for binary data. But at least JSON arrays are a bit more (but not much) memory efficient than XML elements, and JavaScript itself has typed binary arrays, although standard JSON doesn't support them (using base 64 strings sucks). So representing paths as flat strings like SVG, instead of JSON arrays of numbers, is probably still a win for JSON if you're mainly drawing static paths, but not if you want to read, process or edit those paths a lot.
Now, instead of using SVG, it's usually easier and more efficient to just use your own (or some standard) JSON format for structured graphics, and draw it with the canvas api. But there will never be one universal JSON graphics format, just many special purpose ones. For example, d3 has a lot of different JSON formats (domain specific languages) for representing different kinds of graphics and data.
More on PostScript and its history:
https://news.ycombinator.com/item?id=21968175
>Brian Reid wrote about page independence, comparing Interpress' and PostScript's different approaches. Adobe's later voluntary Document Structuring Conventions actually used PostScript comments to make declarations and delimit different parts of the file -- it wasn't actually a part of the PostScript language, while Interpress defined pages as independent so they couldn't possibly affect each other:
And that can be optimised further (by svgo) to:
<path d="M10 10h80v80H10V10"/>
They can get almost hilariously dense and obfuscated.By reusing the SVG format, it is still compatible with existing tools, such as Adobe Illustrator and Inkscape (meaning that they can open SVG Native files).
Adobe even has an open source renderer for SVG Native [2].
Cool, that's only the third time they do that (after Tiny and Basic).
No.
Strong schemas are _why_ it's supported across a multitude of platforms. JSON because "it's got JSON, what plants crave" is an absolutely dumb argument. And the see article "parsing JSON is a minefield". I can't think of a way to make the standard worse than re-implementing in JSON.
There are some legit missing features from SVG as other commenters have pointed out. To me, it sounds like the tooling is actually what's lacking. The equivalent of minifying or even the option to render xml to a binary representation is probably what is actually lacking.
XML Schema is the main reason I use XML. I hate Schema. Hates HAAATESSSS NASSSTY SSSSCHEMA..., but it is ironclad. If it is described in Schema, then it is guaranteed (or not; in which case Schema also accounts for that).
JSON Schema is a non-starter. No one ever seems to use it, and I don't think it ever made it out of committee. It pretty much goes against why JSON is popular.
If anyone wants REAL pain, then we should go back to BNF: https://en.wikipedia.org/wiki/Backus–Naur_form
I worked with a BNF compiler, and I still wake up screaming.
That's work for whoever is producing the document, but wonderful for all the consumers.
I'd argue the same restrictions on "setting handwave to maximum" and "XHTML doc has one missing brace, no web page at all now" are two extremes that have their parallels in the 'dynamic vs static typing', 'immutable pure vs mutable impure functions', 'XML vs JSON (oh wait!)', and 'web app vs native program' arguments that we HN readers love to debate.
The trick is finding the happy medium between both extremes. Training some 'slackers' to "do the job properly" and convincing some 'purists' to be more pragmatic to the limitations of the world around them.
After all, we need something to occupy us during this Eternal September.
And in at least three of the four antagonisms you list, it comes down to the DX of the tooling. Static typing isn't so much of a pain if your type errors are helpful and descriptive. The missing brace in the XHTML doc is easily found when using proper and easy to use XML linters. And I don't see why writing native apps shouldn't be as simple as writing web apps. At least there are attempts that seem to get some mileage out of the idea (e.g. Revery, Flutter).
The trick is to make the toolset more approachable.
The flip side is that you will never find a JSON with the structure "{ "zoom": "startZoomMarkup" }5{ "zoom": "endZoomMarkUp" }".
Often what you want is instead to generate an AST, in that case you will run into ambiguosness problems.
My company heavily uses JSON schema.
The OpenAPI project uses JSON Schemas, and AWS also allows for their usage in places.
And finally, there are a couple good projects that automatically generate JSON Schemas from Typescript definitions, which is my preferred way to go. :)
But note "draft."
Compare to: https://www.w3.org/standards/xml/schema
Hasn't stopped tooling from being built around it.
I fail to see any benefit of a rigid, typed, schema'd JSON that XML does not offer more naturally.
JSON schema let's you make that a bit safer.
The largest benefit of JSON is that is not XML. That it doesn't get bogged down in schema's, strong typing, interfaces, contracts and so on. You loose all that by turning JSON into XML-written-in-a-different-syntax.
Some languages are simple, for example the language that is any sequence of one or more 'A'.
A, AA, AAA, AAAA and so forth are all valid sentences in this language. This is a particularly simple language, but it could be described in BNF:
<S> ::= A | <S> A
This language could also be specified as { A^n | n >= 1 } using a set notation or as a regular expression like A+A slightly more complex language is the language of all sentences made up of a sequence of A and B in any order followed by a sequence of C exactly twice as long as the A's and B's together:
<S> ::= <X>CC | <X><S>CC
<X> ::= A | B
Legal sentences look like: ACC, AACCCC, ABCCCC, BABCCCCCC, etc. Such a language can't be described by classical regular expressions, but can be described by the context-free grammar expressed in BNF above.There are alternatives and generalizations of BNF notations. Wirth used railroad diagrams to describe the syntax of PASCAL in the Pascal Report describing his language. This are more visual, but not more powerful.
BNF might be a pain, but there is generally no simple way to express complex syntax requirements for programming languages. BNF isn't used to solve the same problem being addressed by XML or JSON which are methods to encode data.
But I used to work on an X.400 system, and we would have to map out data structures in BNF, pass it through a BNF/X.409 compiler that always hashed things up (often our bad, but that was back in the "big iron" days, when we would have to arm-wrestle for compiler passes), then, once we fixed the X.409, we'd need to pass it into a X.409/C compiler.
And fix it again...
Ironically, if one wants to implement a SVG reader they have to use the BNFs provided in the spec (for paths and whatnot).
Though the SVG 1.1 DTD is pretty nice to work with, I've built a SVG DOMish generator for both c++ and python (which takes forever to compile) so one, err...me, doesn't have to deal with all the underlying XML craziness. One of these days I'll get around to finishing the path parser, color parser, transform parser, &etc...
XML may be verbose and not readable, but it is an excellent format for describing and storing data.
One alternative for SVG could be an SQL table like datastore, similar to sqlite. I would love yo have the ability to query data, rather than parse through XML. I do understand that transferring such data blobs over networks is an issue.
I agree that use varies widely, but the difference is fairly clear. The existence of nonsense doesn't mean one can't make good decisions.
to use anything else means relying on an arbitrarily-complex custom schema rather than the much-simpler-and-far-more-standard schema of the language itself. similar to how we use datatypes in programming languages rather than just strings everywhere. both work! one is clearly giving up a lot, and gaining little.
I don't know if you can slim down vector graphics successfully. Every SVG renderer in existence slims down SVG to some degree. The ones that do it the most are used the least because they can't sufficiently express the intent of the designer.
Surely SVG does not include months of work which is entirely useless. What features are we giving up for this simplicity?
Even just a linearization of most of SVG into the obvious layout, simply dropping any problematic features (script, XLink, namespace extensions) would probably be a pretty good start that you could churn out in a week as a prototype. You'd get an efficient binary format, an obvious mechanism for translating existing SVG into it (with equally obvious mechanisms for warning users about the existence of untranslated features in the output; also an easy mechanism for scanning sample SVGs and getting a statistical idea of how many fit into your new scheme), an obvious mechanism for translating your format back into SVG, and serialization support as cross-platform as your choice of serialization mechanism right out of the box. If this is your goal, this has a pretty hard to beat bang-for-the-buck.
Or too much abstraction and extensibility leading to excessive complexity!
I don't know if any of the existing Protobuf-like things have these built in, though. I know some of them have variable integers in them but I don't know about variable floats.
This is kind of what I was getting at in that we have a lot of prior art now. Someone putting variable ints or floats into SVG no longer has to create their own bespoke format, and proceed to fall into the various traps themselves and accidentally write it forever into the spec; there's a lot more of these sorts of pieces that you can pick up off the shelf than there used to be.
Am I missing something, or can I not read 2 bytes from a file and detect if that's the end of the number or I need to read the next 2 bytes?
for this purpose it would be enough to revert the order of the stream.
You don’t even have to use byte-sized fields, but doing that makes encoding and decoding harder, and the overhead of also storing the actual lengths in bits of variable sized fields may be too much to make it worth that.
Well unless it doesn't. I've recently read more than one blog post that claimed JSON was actually borderline faster—sometimes more than borderline—than some popular binary formats like protobufs. I won't link those benchmarks here because none were well enough conducted or presented to warrant a link; I myself treat those reports as plausible but anecdotal hints.
It's not like saying "binary format" makes all your troubles go poof. You still have to parse it. Yes, you can control it better when you create it, but when you've created it and it sees a lot of adoption, then updating it will be just as troublesome as updating any popular format, meaning many clients won't support some of the new features for at least some time.
Numbers of course are central in any actual SVG, and it turns out using a textual format to send numbers is not as stupid as it sounds, space-wise: `1` is just one byte, but it's 8 bytes (I think) of RAM when turned into a JS IEEE754 number. You can represent arbitrary precision or restrict yourself to integers on a whim when using text, something probably not true of most binary formats.
One of my qualms with SVG are that XML itself is a needlessly complex format. As sort-of-proof for that claim I present you the fact that the [namespace on the `href` attribute has been elided from the standard](https://developer.mozilla.org/en-US/docs/Web/SVG/Element/use). This of course is just the tip of the iceberg of XML complexity.
The other big Qualm is that for some reason the authors of SVG felt it necessary to open the doors for a lot of variation on how to write things like e.g. the `<path>` element's `d` attribute, like you can add commas or leave them out, you don't need whitespace in `"45-67"` because that can be parsed as `(45,-67)` and so on. Not useless stuff, but in the end more than minimal complexity.
I find text-based formats so much better in many regards. Just look at the rich appearance of web pages today, almost all of that has been done using text-based formats. They're really powerful, and where needed, may be compressed using a generic algorithm.
The thing is all these languages (ASM, C, XML, JSON, Python, etc) are fundamentally about communication between a human and a machine. We make tradeoffs between making them easy to parse (for the machine) or easy to read and write (for the human) but they all must address both needs. And when you have a problem domain as large as "2D Vector Graphics" or worse, "The Web", complexity is unavoidable.
I'm not saying we can't do better than SVG. I agree with many of the author's points and would love to see a minimal subset of SVG that is easy to write by hand. The tradeoffs are engineering decisions and important. But don't pretend that you can pick a goal such as, "easy for humans" and ignore all the other dimensions of the language design problem.
True! And we should not postulate who is at both sides of the communication - machine or human. It's like postulating that a hammer is a tool which must be held by a human hand, which then precludes robot-operated hammers.
I'm feeling equal parts: "this is beautiful", "this is perverse", and "I want to do this too"
The one I was proudest of was a 3d chemical viewer where you could grab the molecule and spin it. SVG was only used because it was convenient to do the visualization. All of the 3d transforms were done in Javascript. But to have Javascript that could fluidly spin a 3d model was pretty amazing in 2000.
https://gist.githack.com/mejedi/90eac4f11f794f43808f90e38b1a...
https://upload.wikimedia.org/wikipedia/commons/6/6c/Morphing...
A morphing digital clock
FYI, the OP posted an update with a reply to you:
> This article was posted on Hacker News... > In the comments there somebody mentioned an interesting use of the <script> tag in SVG. I’m unsure though whether I should be impressed or horrified :D
"Why am I still struggling with browsers having different SVG implementations 18 years after the initial implementation?" is something I had to ask my self two days ago when the comma in our logo magically shifted to the left only on Firefox.
It looks like the spec here is partially to blame because 800 pages is something that even the most meticulous team of devs would have a hard time implementing without some divergence from other implementations.
I decided to switch to a high quality PNG and call it a day since we had other more important things to be focusing on.
Dreamweaver used to do this for HTML, sometime before it got bought by Adobe. Today Dreamweaver seems like it would be usable on a 90 THz computer but I can do 10 push ups waiting for the UI to settle down after typing "Hello World".
A split-screen SVG editor would be particularly good, since so much of editing an SVG file is getting the numbers right. You'd face many interesting problems, such as exposing decimal math to the user (whatever iffy floating point is used, your git commits are going to be awful unless you "snap to grid" in factors of 1/2 and 1/5.)
What I'd particularly like to do is have control of the structure of the document so I can group elements with the discipline of a programmer.
A tool like that still might not be satisfying because Bézier curves were made to make people feel like they were bad at computer graphics and took over instead of better alternatives that were patented at the time. I was shocked to find that my favorite character designers don't know how to draw anime characters with Bézier curves.
Works pretty well as split screen editor. The preview refreshes on save with ctrl-S.
"The <g> SVG element is a container used to group other SVG elements.
Transformations applied to the <g> element are performed on its child elements, and its attributes are inherited by its children. It can also group multiple elements to be referenced later with the <use> element."
That is the most prominent tool one would use to apply "programming" or "architectural" discipline to SVG.
For instance you could make a <g> that looks like a cat paw and then <use> it to make a track of cat paws. That's more disciplined then say, "cut-n-paste" a cat paw you drew 10 times which a real artist might do, or I might do because I can always figure out how to do it with a drawing program.
Related to that is the use of CSS, which can range from beautiful code which is easy to understand towards a complete mess.
You definitely don't want to solve SVG's bloat problem by moving to another text-based format.
The CE spec is for all intents and purposes done (I don't anticipate any more changes), and the reference implementation is in the final stages. After that comes the schema format, then the official V1 release, and then I can move on to technologies built upon it.
From some back-of-the-envelope calculations I made recently, I should be able to represent the same vector graphics in the same basic structure (such that you could convert between them), but in about 1/3 the space (not including the crazy stuff like CSS and JS).
http://slides.com/sdrasner/svg-can-do-that
Sarah Drasner is awesome, and even though the preso is 4 years old(!) it's still relevant and inspiring.
1. It doesn't support premultiplied alpha interpolation mode. If you ever import partially transparent rasterized images then it can cause visible artifacts at the edges of opaque regions.
2. Last time I checked no implementation supported interpolation in linear colorspace (linearRGB [1]).
So even though it's bloated it still misses features. And even though browsers are good fit for supporting it, they still don't bother implementing the full spec. And they implement PDF with vastly more gradient elements and color interpolation modes.
[1] https://www.w3.org/TR/SVG11/painting.html#ColorInterpolation...
- Relative coordinates. You can hack around this by using other SVG documents inside an SVG document, but this is really just a hack.
- Conic gradients. Even CSS has them.
Also, there are still a few implementation bugs for SVG. The Chromiums and Firefox teams have done an outstanding job at fixing bugs over the last few years, but if you do advanceed stuff, you will inevitable bounce into one of them. Also, Safari is not really good at SVG, but honestly, who cares.
You can also use <def/> to achieve something like that.
AI4 stands for Adobe Illustrator 4.0. Gnuplot and Mathematica can produce AI4 files. Most vector graphics programs can open AI4 files too.
Here's the documentation: http://www.idea2ic.com/File_Formats/Adobe%20Illustrator%20Fi...
The postscript language reference manual is 912 pages.
https://www.adobe.com/jp/print/postscript/pdfs/PLRM.pdf
The PDF reference is 978 pages.
https://www.adobe.com/content/dam/acom/en/devnet/pdf/pdfs/pd...
SVG, if it's less than either is actually doing well...
IMO these three technologies and the accompanying docs are masterpieces, despite bloat. Digging in to those reference docs can almost be fun
The author's idea of SlimSVG sounds interesting, but I think focusing on the file format isn't enough -- one powerful feature of SVG is that it's part of the DOM. Can there be a better system than the DOM+JS for allowing a dynamic document?
EG, document object model in JSON (maybe with JSON schema) and allow updates with JSON patch, with a client/server model for updates. (There's more to DOM)
Personally I think XML is well suited to this task, and am glad to have learned of SVG Native from this thread
> Instead you compile it into an PDF (which is also a horrible format and badly bloated, but well)
Yeah... PDF might actually be worse in some ways (and it can have javascript too). With SVG, it isn't too hard to be able to use a subset of SVG to generate images. With PDF though, just making a valid, useful PDF is pretty complicated and requires a decent understanding of the spec (which is big). Being able to properly render any valid PDF is probably even worse.
If we're going to start over, let's create a new syntax that's purpose-built from the ground up. JSON doesn't replicate all of the drawbacks of XML for this kind of purpose, but it replicates a lot of them. Like, all the forces that lead major XML formats to embed their own little mini-DSLs inside of strings are present in JSON as well.
There's this tradition in our profession, perhaps born of a reaction to the "not invented here" syndrome of the '90s, to habitually hack things together out of bits and pieces that already exist. Seemingly for no reason other than because they already exist, and they can be hacked. I've recently started calling it "Wallace and Gromit engineering."
https://progenygenealogy.com/Portals/0/images/charts/Kennedy...
and Fan charts:
Ghostscript does not have a SVG output driver, although you could first output to PDF, and then use another program to convert PDF to SVG. If they make SlimSVG or other vector formats, then hopefully a Ghostscript driver for that format could be provided, in order that it would be possible to produce it from PostScript programs.
You could then produce SlimSVG or whatever other format using whatever programming languages you wanted, whether PostScript, Haskell, or something else.
Which is a shame, because I actually agree with a lot of it.
I propose web standards be refactored and split into 3 categories: A) Art, media, & entertainment. B) Documents and brochures, and C) CRUD, GUI, Data, productivity
They would overlap as much as practically possible, but could also specialize in their respective niche better.
Concerning the simple use of standard shapes, you could restrict yourself to simple shapes within SVG : thus youd already have myriads of existing implementations and tooling, it would be mostly understandable by humans, short and writable by hand/simple scripts.
You could even call them mobile SVG Tiny and Basic : it exists since 2009. https://www.w3.org/TR/2003/REC-SVGMobile-20030114/ . That would be a good start.
Man oh man what a rabbit hole. I went back and forth on implementing mostly the GUI with CSS to mostly the GUI with SVG and now BACK to having mostly CSS as opposed to SVG's.
Biggest problem is the blurriness ! "Scaling a playing card with text down (transform:scale|scale3d etc... ) - just ends up looking blurry. These are vanilla html-based-svg I 'coded' by hand.
Just for the note, there is such format already and it is quite popular actually, see: https://lottiefiles.com/
Lottie files are JSONs produced by Bodymovin plugin for Adobe After Effects.
These are primarily for short vector animations but lotties work OK for static images too.
SVGs may not be perfect, but the pros definitely outweigh the cons. When I work on the frontend, I think there are very few use cases where I have to worry about the readability of an SVG. The author also kind of lost me at JSON-based graphics.
IconVG is a compact, binary format for simple vector graphics: icons, logos, glyphs and emoji.
IconVG is similar in concept to SVG (Scalable Vector Graphics) but much simpler. Compared to "SVG Tiny", it does not have features for text, multimedia, interactivity, linking, scripting, animation, XSLT, DOM, combination with raster graphics such as JPEG formatted textures, etc.
The announcement is at https://groups.google.com/forum/#!topic/golang-nuts/LciLeI5p... and quoting from that:
"The SVG spec prints as 400 pages, not including the XML, CSS or XSLT specifications. [The IconVG godoc.org page, which includes the file format specification, prints as 26 pages]...
The Material Design icon set... consists of 961 vector icons. As SVG files, this totals 312,887 bytes. As 24 * 24 pixel PNGs, 190,840 bytes. As 48 * 48 pixel PNGs, 318,219 bytes. As IconVGs, 122,012 bytes."
I'm honestly horrified that someone out there is writing SVG by hand.
I first did it with inscape and the file size of the simple logo I was editing went up by 6x due to the inscape metadata so I did it manually.
But yeah, I did. Quite a few times. I even did a SVG path to G code converter once.
Inkscape made a nice initial template, but lacked finesse and precision with naming, scaling, text-wrapping and nesting. Inkscape came up with a rather messy blurp of XML. No clear IDS, rounding errors, arbitrary nesting through groups, arbitrary ordering, etc.
"SVGO's Missing GUI". Give it to the designers and they can make sure to give you the smallest possible file that makes sense. I find that some files look great at "0" precision, some require "1" or "2".
I'd prefer to have stuck with SVG had it not been for the I'd collision thing, so with one of these tools in our chain we just might do that.
We use some server side code to recolor svg's on the fly with caching headers so they aren't requested all the time, mostly through img tags since we moved from PNG to SVG, its just works.
Using inline SVGs also reduces our caching complexity _dramatically_. I'm now "just" caching markup with no binary asset caching.
Moving forward I'm going to experiment with
At least SVG as images can be resized so you only need one asset regardless of size, color manipulation then is the issue, but since SVG is just xml it pretty easy to do server side and again the browser will handle caching just like any other image.
So we jumped to PNGs (like 2 weeks ago, fyi). Now that we (know about, and) can use one of the svg cleaners, we'll switch to that and go back to inline SVGs.
PNGs were always seen as inferior. We don't want extra requests to load them - our app is very sensitive to the number of requests we make - we're importing them as datastrings even.
Anyway, the choice was made. Now we're 90% likely to revert to inlining SVGs with some kind of cleaner in the pipeline. Please don't lose any sleep over it. We'll be just fine.
Ah yes, good ol' https://xkcd.com/927/