Typed-Html: Type Checked JSX for Rust
github.com
github.com
Basically, with Rust, you can import syntax like this much like you'd import a library in other languages.
Macros aren't complex, c preprocessor macros are complex. Although generally, magic that introduces names that aren't defined anywhere is bad, whether it's c macros or a python metaclass, but the metaclass at least has some direct connection to the magic name.
At its core the C preprocessor language is a simple token substitution scheme (meaning it strictly operates on the word level). I couldn't write one without looking at the spec because it has some ugly edges. But fundamentally it's a simple technology. It's flexible but that also means it must be used carefully.
[0]: http://www.semanticdesigns.com/Products/DMS/DMSToolkit.html
I'd say the main cause of complexity with C preprocessor is that it indiscriminately feeds it's output to a complex system. This means that relative small perturbations such as precence or lack of parentheses can have large and dramatic effects. One could say that a system that works around these by adding safeguards such as hygiene-by-default is more complex by it self, but it definitely begets more controlled behaviour.
Abstraction is the act of hiding necessary, but irrelevant, complexity.
Timsort is much more complex than quicksort or merge sort. A basic doesn't need to know that they're using timsort instead of mergesort. An advanced user may prefer knowing that timsort will always be faster, but not knowing this doesn't harm them.
The same goes for many (most?) performance optimizations. A more complex implementation is often wholly transparent to the API (although I guess you could argue that the abstraction is leaky since the system runs faster), but simultaneously often warranted when applied across all users of a system.
The same thing is true with macros that act on an AST instead of raw text. You use a wholly different API (acting on a tree structure instead of a string), but this makes many changes easier, and makes them all safer, at the cost of the API having to handle parsing the language for you.
There's no downside to the end user. Only a higher cost of implementation. You're free to disagree, but you should be able to demonstrate specific examples of costs to the end user in those two cases, as well as any others.
There's a nice quote, "Being abstract is something profoundly different from being vague. The purpose of abstraction is not to be vague, but to create a new semantic level in which one can be absolutely precise."
If the user really does not care what implementation is used, you could use any. You could use a less sophisticated one. In any case I do not think I had a sort implementation in mind, since one might argue it's not really complex. It doesn't interact with the rest of the system in complex ways. If that makes sense.
UPDATE: Now I know what bugged me in the way you put it initially -- it's in "complexity of the implementation leads to a simpler abstraction/conceptual model/end user experience". Note that this does not apply to Timsort (or any other sort). The "complex" implementation does not make for a simpler conceptual end user experience. A sort is a sort.
I don't have anything against AST transformations. They are a good idea to implement if one can figure out relevant usecases and a usable API. But in most cases I guess personally I'm likely to prefer either that 1-line macro, or not to add tricky AST manipulation code to create heavy syntax magic for a one-off thing.
And under this definition, using AST level macros is vastly superior to text-transforming ones. Text-transforming macros don't allow the user to be precise. Or at least, to be precise, the user must know a number of rules that are arcane and not obvious. `#define min(X, Y) (X < Y ? X : Y)` feels like it should work, but it doesn't. Or well it does sometimes, but not always.
I'll come back to this in a second, but lets talk about sorting. You claim that if the user doesn't really care you can use a less sophisticated abstraction. But that's not at all true! In fact, its a violation of the Liskov Substitution Principle[0]. If I have an interface `Sort`, I shouldn't care whether it is of type Bubble, Tim, or Quick. But if I have a TimSort, I certainly may care if you swap it out with BubbleSort. In this case, the desirable property is speed. One way of putting this is that for any Sorts S and T, S can only be a Subtype of T if S is faster than T. This allows you to replace a T with an S, and you won't lose any desirable properties[1].
Another way to put this rule would be that when modifying an API, your changes shouldn't cost the user anything. Replacing Timsort with bubblesort costs the user speed. That's bad. It violates an assumption the user may have had.
And while you say that this doesn't make for a simpler conceptual end user experience, I disagree. If the system just is fast, you don't need to worry about performance. If `sort` is fast, you won't need to implement your own faster sorting function (or even have to try and figure out that that's what you need). Not having to write an efficient sorting algorithm certainly sounds simple to me!
Similarly, there is a cognitive cost to an API. And working with an API that is astonishing[2] has a cost. Its harder to reason about how it will interact with the rest of the system. An API that can guarantee that all macro results will be syntactically valid is simpler than one that requires that you manually do that bookkeeping. Same with a language that can guarantee that all memory accesses will be valid. There may be associated costs with these things (you can't string-concat syntactically invalid strs into valid ones, or you can't implement certain graph structures without jumping through hoops), so its not quite as clear cut as with sorts, but I think its pretty clear that AST-based macros are simper to interact with than Cs.
In fact, C's macros being so simple makes them more dangerous. Its easy to understand how they work internally, so a novice may feel like they won't be surprising (until...). But the simplicity of the implementation leads to footguns when they interact with other things. The abstraction C macros provide is conceptually simple but leaky, or to use your words, imprecise. You have to understand them more than you think to be able to interact with them safely.
AST based macros on the other hand aren't leaky. They're more difficult to conceptualize, but you don't really need to fully conceptualize, because they won't surprise you. You take in some expressions and modify them, and you'll get out what you expect.
Doing AST based transforms and substitutions instead of text-based ones significantly reduces the cognitive overhead. You stop having to worry about the edge cases where some transformation might happen in the wrong context, or not happen in the right one (as a simple example, applying substitutions via regex vs. via AST means that you no longer have to worry about where there were line breaks).
>I don't have anything against AST transformations. They are a good idea to implement if one can figure out relevant usecases and a usable API. But in most cases I guess personally I'm likely to prefer either that 1-line macro, or not to add tricky AST manipulation code to create heavy syntax magic for a one-off thing.
I'm a bit confused here. I'm not saying that the macros are harder to implement for the end user, in fact the opposite. But that they're more work for the language to implement. A text based macro like
DEFINE_handler(type, function_body) (void handle##(type)((type) input) { (function_body) };
isn't significantly easier to define than something like DEFINE_HANDLER(type t, AST function_body) {
f = ast.Function()
f.name = "DEFINE" + t.name
f.body = function_body
return f
}
In fact, its arguably clearer what's going on in the second example.[0]: https://en.wikipedia.org/wiki/Liskov_substitution_principle
[1]: Yes I realize there are other desirable properties of sorts, such as stability and inplaceness but I'm simplifying.
[2]: https://en.wikipedia.org/wiki/Principle_of_least_astonishmen...
#define LENGTH(a) ((int) (sizeof (a) / sizeof (a)[0]))
#define SORT(a, n, cmp) sort_array((a), (n), sizeof *(a), (cmp))
#define CLEAR(x) clear_mem(&(x), sizeof (x))
#define BUF_INIT(buf, alloc) \
_buf_init((void**)(buf), (alloc), sizeof **(buf), __FILE__, __LINE__);
#define BUF_EXIT(buf, alloc) \
_buf_exit((void**)(buf), (alloc), sizeof **(buf), __FILE__, __LINE__);
#define BUF_RESERVE(buf, alloc, cnt) \
_buf_reserve((void**)(buf), (alloc), (cnt), sizeof **(buf), 0, \
__FILE__, __LINE__);
#define RESIZE_GLOBAL_BUFFER(bufname, nelems) \
_resize_global_buffer(BUFFER_##bufname, (nelems), 0)
#define MSG(lvl, fmt, ...) _msg(__FILE__, __LINE__, (lvl), (fmt), ##__VA_ARGS__)
#define FATAL(fmt, ...) _fatal(__FILE__, __LINE__, (fmt), ##__VA_ARGS__)
#define UNHANDLED_CASE() FATAL("Unhandled case!\n");
#define ABORT() _abort()
#define DEBUG(...) do { \
if (doDebug) \
_msg(__FILE__, __LINE__, "DEBUG", __VA_ARGS__); \
} while (0)
And a clever one, saving a lot of typing (which many will argue only fixes a problem of C itself. But still). #ifdef DATA_IMPL
#define DATA
#else
#define DATA extern
#endif
DATA char *lexbuf;
DATA char *strbuf;
DATA struct StringInfo *stringInfo;
...
None of these is a maintenance burden, and each makes my life significantly easier. I don't believe there's a different scheme that is a better fit here.In many languages, those macros aren't things you'd ever need to do. You're just being forced to make up for a flaw in the platform.
I'm not sure what you mean here. ast-macros can still wrap ast-macros.
And yes, I'd absolutely claim that not tracking array size at compile time is a flaw in C (rust fixes this, you can pass `&[int]` to a function (a reference to a compile-time-fixed-size array) and call `.len()` on the argument. This has no runtime cost in either space or speed).
In the same way that I talk about cognitive overhead above, the requirement that a user manually pass around compile time information is dumb. Note that in C this wouldn't have prevented you from down-casting an array to a pointer, its just that the language wouldn't have forced this on you at every function boundary.
The only reason C didn't do this is because the implementation was costly to the authors. It doesn't have any negative impacts to the end user (well, there's an argument to be made that there was a cost to the end user at the time, but I'm not sure how much I believe that).
Yes, although sometimes text macros are useful, but ast-macros are generally much better yes.
>In the same way that I talk about cognitive overhead above, the requirement that a user manually pass around compile time information is dumb.
If the macro facility is sufficient, it would be implemented by the use of a macro; you do not need to then manually write it during each time. In C, you can use sizeof. Also sometimes you want to pass the array with a smaller length than its actual length (possibly at an offset, too).
- OOP with garbage collection: Need to pass around containers and iterators by references. Bad for modularity (the container type is a hell of a lot more of a dependency than a pointer to the element type). And not everybody wants GC to begin with, not in the space where C is an interesting option.
- Passing around slices / fat pointers with size information. Not as bad for modularity, but breaks if the underlying range changes.
- Passing around non-GCed pointers to containers (say std::vector<int>& vec): Again, more dependencies (C++ compilation times...). And it still breaks if the container is itself part of a container that might be moved (say std::vector<std::vector<int>>.
- With Rust there is now a variation which brings more safety to the third option (borrow checker). I don't have experience with it, but as I gather it's not a perfect solution since people are still trying to improve on the scheme (because too many good programs are rejected, in other words maybe it's not yet flexible enough?). So it's still unclear to me if that's a good tradeoff.
None of these options are orthogonal language features, and #2 and #3 easily break, while the first one is often not an option for performance reasons. All are significantly worse where modularity is important (!!!).
I personally prefer to pass size information manually, and a few little macros can make my life easier. It causes almost no problems, and I can develop software so much more easily this way. I have grown used to a particular style where I use lots of global data and make judicious use of void pointers. It's very flexible and modular and I have only few problems with it. YMMV.
The situations where the borrow checker can't work are different than those that involve arrays. You don't lose anything.
There doesn't always need to be a trade off.
Of course it is. The first is accessible by everybody who knows the host language's syntax, the second requires an understanding of how the syntax is mapped to the AST, which may be not be 1:1 in certain edge cases.
Getting a struct pointer and modifying it would be much more familiar to anyone who hadn't yet written macros than jamming together tokens with ##.
I also don't like putting the pointer together with the extra data because that means dependencies, and also I want a simple bare pointer to the array data (instead of a struct containing such pointer) that I can index in a simple way.
I also don't like the stretchy_buffer approach where you put the metadata in front of the actual buffer data. Again, because of dependencies.
The alternative would have been to go for C++ and make a complicated templated class. I don't use C++ and templates are a rathole on their own. So a single ## in my code solves this issue. I'm happy with it for now.
I'm suggesting that in lisp or rust macro land, your macro is a lisp or rust function. So in sane c macro land, your macro is a c function.
macro macrofun(AST* node) {
node->name = strcat("BUFFER_", node->name);
}
Literally just use a subset of c syntax on a c ast represented as a tree of node structs. It need not be blazing fast, it runs at compile time.C++ constexpr is close, although decidedly less powerful, but is still a huge win over c macros.
You keep making excuses for why you're doing these things, and I don't really care why you're doing them. I'm saying you shouldn't need to, and that the interface with which you would solve them should be less awful and inherently error prone.
But making a less error prone interface takes more up front complexity. You've elsewhere claimed that this has positive externalities ('it forces you to understand c sytax better'), but I'd reverse that and claim that
1. It violates a common api ("languages are expression oriented"), and is therefore both astonishing and leaky
2. It doesn't force you to understand the language better, it prevents you from being productive until you understand a set of arbitrary rules (you list these elsewhere) that aren't necessary for normal use. Token based macros require you, the user, to understand how c is parsed, AST macros don't because they do it for you.
In the end, there are zero problems with my approach here, either. And the CPP doesn't encourage you to get too fancy with metaprogramming. Metaprogramming is a problem in its own because it's hard to debug. I've heard more than one horror story about unintelligible LISP macros...
Note that I am going to experiment with direct access to the internal compiler data structures for my own experimental programming language as well. But you need to realize that this approach has an awful lot more complexity. You need to offer a stable API which is a serious dependency. You need to offer a different API for each little use case (instead of just a token stream processor). If you're serious like Rust you also need to make sure that the user doesn't go all crazy using the API. Finally, it's simply not true that you need to understand less about parsing (the syntax) with an AST macro approach. The AST is all about the syntax, after all.
Not knowing whether `f()` is a syntax error or not removes a huge part of the foundation that people use to read code they do not know 100%
And this is exactly why they are useful. You can do things with it that you cannot do in any other way. Practical things, I want to mention.
Some examples: conditional compilation based on feature set. Automatic insertion of additional arguments (like __FILE__, __LINE__, or the sizeof of arguments). Conversion of identifiers to C strings. Computations such as getting the offset of a member in a struct. Ad-hoc code and datastructure generation.
Many of these could be individually replaced with complex features built into the core of the language, by making arbitrary ad-hoc decisions that would be hard to standardize and would probably kill the language.
> Not knowing whether `f()` is a syntax error or not removes a huge part of the foundation that people use to read code they do not know 100%
It's your responsibility to match parentheses when required (which is almost always). That's easy for a macro that is usually 1, or at most 5 lines. And if something goes wrong, it's usually not that hard to find the error and fix it. You need to be aware of some gotchas though (wrap expressions in parens, don't do naked if statements or blocks but wrap them in do..while(0)). That means you will get to learn the C syntax more intimately.
This is completely disregarding C and its token macros.
These end up existing anyways though. At least a few places are using Go with their own preprocessor - isn't Kubernetes doing this?
For example, how do you step a macro expansion in Rust at any given point in a source level debugger?
Yep, you don't.
Even on Lisp, if we want something more developer friendly than (macroexpand) the only option are the commercial Common Lisp environments.
Having said this, I also do like having macro support around, provided it gets used judiciously.
> Yep, you don't.
I don't see any technical reason why that can't be done
The point it that such kind of tooling is generally not available in FOSS offerings.
Yes it can be done, but it isn't currently available.
Any plans to make it work on stable?
We'd of course like to have it on stable, like everything, there's just more pressing stuff to stabilize.
https://blogs.msdn.microsoft.com/vcblog/2018/05/07/macro-exp...
Or you convert them to C++ constexpr templates, which you can easily step on the debugger.
But I have to ask, will the choice of the MPL as the license have any effect on its potential adoption? Yew is Apache/MIT, matching Rust and much of the ecosystem, for context.
MPL, for folks who aren't aware, is a successor to the CDDL seen in Solaris etc. and is a per-file copyleft, thus having a nicer legal structure while getting the same rough benefits of GPL or at least LGPL. The idea is “any modifications to this file must also be open-sourced under MPL, but you can package this file with proprietary other files that are not MPL and integrate them into a larger proprietary thing as long as your modifications to this file alone are open-sourced.” The goal is to protect weakly from Microsoft-esque “embrace, extend, extinguish” as GPL does but enable commercial integration the way BSD does.
In practice nobody seems to be all that pissed at Mozilla because of their license; the MPLv2 added GPL compatibility so the GPLers are mostly able to use MPL software and it gives a nod towards the stronger copyleft they like; commercial applications which just use the software as a library don't mind.
I'm very glad that this is happening, the whole project got started when I was struggling to write PHP/JS in the early 2000s and the itch was getting worse and worse.
The key idea is to traverse a tree datastructure which will map to database entities, for instance:
database my_db {
int /basic/i
float /basic/f
string /basic/s
stored /r
intmap(stored) /imap
}
Kind of an ORM but without objects (Opa is purely functional) and without relations when using NoSQL.But DbGen has support for several types of database including PostgreSQL: See for instance this blog http://www.josetteorama.com/to-sql-from-nosql/
<!-- The table element
Content: In this order:
optionally a caption
element, followed by zero
or more colgroup
elements, followed
optionally by a thead
element, followed by
either zero or more tbody
elements or one or more
tr elements, followed
optionally by a tfoot
element, optionally
intermixed with one or
more script-supporting
elements. -->
<!ELEMENT table - -
(caption?,colgroup*,
thead,
(tbody*|tr+),tfoot?)
+(%scripting;)>
Regular content types are pretty fundamental to markup languages.> The structure validation is simplistic by necessity, as it defers to the type system: a few elements will have one or more required children, and any element which accepts children will have a restriction on the type of the children, usually a broad group as defined by the HTML spec. Many elements have restrictions on children of children, or require a particular ordering of optional elements, which isn't currently validated.
It's not complete or thorough at present.
Regarding your link: I'm impressed someone took the time to write a DTD for HTML5!
Haskell has had something similar for many years now [1]. Ultimately the development cycle can become a big issue for this kind of thing (how long does it take to go from editing your type-checked template to being able to see that in the browser). In the Haskell version we did have a development version that allowed certain types of changes to show up quickly.
[1] https://www.yesodweb.com/book/shakespearean-templates#shakes...
https://ocsigen.org/eliom/1.3.4/manual/html
The syntax is fairly lightweight, just << followed by valid, type checked XHTML ended by >>. There are also antiquotations so you can use it as a templating language.
For those of us generating XML from C I wrote a hairy set of C macros:
https://github.com/libguestfs/libguestfs/blob/master/common/...
so you can write code like this (which is not fully checked at compile time of course):
https://github.com/libguestfs/libguestfs/blob/4aa712d551f9d4...
Note that tyxml goes quite further than Rust's typed-html: the nesting is significantly more flexible, type inference is still complete, and it will verify additional properties like "don't use <a> inside <a>". It can also be used conjointly with reactive and/or isomorphic programming.
[1]: https://ocsigen.org/tyxml/ [2]: https://ocsigen.org/tyxml/4.3.0/manual/ppx
- How compositional it is ? I have find that some of the HTML properties are very hard to verify in a compositional way, see https://github.com/ocsigen/tyxml/issues/175
- How do you handle the "subtyping" aspect that is intrinsic to HTML ? Or phrased in another way: what's your type encoding ? :)
- I suppose you desugar to a set of combinators, but you don't really expose those. Why ?
Why indeed. I personally find these "HTML in a programming language" tools pretty unappealing, because you have to give up all the normal tools of the programming language while using them [1].
Whereas you can usually write a really simple internal DSL to describe HTML that is as clear and concise as HTML, while being as convenient and powerful as the programming language. My quick and dirty efforts at this usually end up looking something like this:
html(
head(
title("My First Page")),
body(
h1("Welcome"),
p($("Made by "), a(href(), $("Tom"))),
userIsOldSchool()
? p("Bring back HTML 3.2!")
: div("Made with HTML 5.")))
The typed-html crate has an object model like this underneath, but doesn't, AFAICS, expose it.[1] With the honourable exception of ScalaTags, which i still wouldn't touch with a bargepole
A component example would be very useful. Being able to abstract away chunks of your views into discrete parts is a pretty important feature.
One big thing that's not mentioned in the docs is whether or not it handles escaping values for you, as well as an equivalent of React's dangerouslySetInnerHTML [0] when you wish to embed HTML.
[0] https://reactjs.org/docs/dom-elements.html#dangerouslysetinn...
http://www.typescriptlang.org/play/#src=%0D%0A%0D%0Atype%20A...
[1] https://docs.rs/typed-html/0.1.0/typed_html/elements/struct.... [2] https://docs.rs/typed-html/0.1.0/typed_html/elements/struct....
let doc: DOMTree<String> = html!{
<body>
<img alt="xxx" />
<p alt="xxx"></p>
</body>
};
If you try to give <p> an alt attribute, the program fails to compile. error[E0609]: no field `alt` on type `typed_html::elements::Attrs_p`
--> src/main.rs:12:16
|
12 | <p alt="xxx"></p>
| ^^^ unknown field
|
= note: available fields are: `accesskey`, `autocapitalize`, `class`, `contenteditable`, `contextmenu` ... and 9 othersI think that the best solution in Haskell would be to group the types in a typeclass. However then the typeclass is not directly associated with the p or img value constructors. I wonder if there is some way to use dependent types to improve this situation.
<p>hello</p><img src=#>
If you try to access the `.alt` property of the <p> tag you get `undefined`, and if you try to access the `.alt` property of the <img> tag it returns the empty string in this case because the property exists on that tag, but hasn't been set to anything.And I imagine its quite a bit more work than just the tags
Render to a virtual DOM... pass it on to your favourite virtual DOM system
Curious what'd it be like to wire this up with virtual-dom or maybe even React itself.Also curious how this compares to https://github.com/DenisKolodin/yew (which I'm not familiar with).
Edit: Sure, having automatic HTML escape for values is good. I work with JSX every day, and I can't really imagine having it all in one huge chunk and no code reuse. Also, what about conditional hide/show and lists?
Edit: You could probably create a separate function with an `html!` return and call that from the first? If the return types line up, which they should, I don't see why you couldn't do that.
<table>
<caption>Example Table
<thead>
<tr>
<th>Col 1
<th>Col 2
<th>Col 3
<tbody>
<tr>
<td>One
<td>Two
<td>Three
<tfoot>
<tr>
<td>Four
<td>Five
<td>Six
</table>
I'm currently drafting an article of HTML you never need to write, currently I'm covering:- some start tags are implied and never need to be written unless you're adding attributes
- some end tags are implied and never need to be written
- default attributes never need to be specified
- element values often don't require quoting
- no trailing slashes are required on void tags
Wow.
HTML in Rust? No!
There is no pseudo javascript involved here.