A Prettier JavaScript Formatter
jlongster.com
jlongster.com
[0]: http://homepages.inf.ed.ac.uk/wadler/papers/prettier/prettie...
Even if you can improve "prettiness" with heroic technical efforts, I think the end result would be uglier. There's something special about simple rules coming together to address just the most important needs.
In the tech world, we tend to put more finesse into things... because we can. But as a result, most software is over-finessed and therefore not finessed at all. A great example is how tortured we let our CSS rules get, rather than allowing a design to relax a little into the constraints of the layout engine.
A hand-set newspaper isn't beautiful because it overcame every constraint, it's beautiful because it accepted its constraints, and made hard tradeoffs in service of a goal.
Bob Nystrom (munificent) disagrees[0] after writing one himself:
> The search space we have to cover is exponentially large, and even ranking different solutions is a subtle problem.
[0]: http://journal.stuffwithstuff.com/2015/09/08/the-hardest-pro...
My (mostly self-imposed) task was more difficult because I was trying to follow the existing style that humans were hand-applying to their Dart code, and that had a lot of tricky non-local cases that look nice but are hard to automate.
[0] https://github.com/dart-lang/dart_style/wiki/FAQ#why-cant-i-...
> I'm never afraid to peek into unfamiliar code because of that.
This is exactly why dartfmt exists. It's not about making your code more readable to you, it's about making strangers' code more readable, because that lowers the bar to contribution between people in the ecosystem.
Also, describing how these building blocks interact using laws is really incredibly useful! For that, credit goes (I think) to Hughes [1].
Having choices can be nice, but it introduces a huge amount of cognitive load. And, really, if we all just use the same format, we'll all (eventually) get used to it.
We really do have better things to deal with than formatting style.
(By the way, for people trying to do the same, I recommend the redbaron library, which makes it much easier to handle: https://github.com/PyCQA/redbaron)
gofmt's decision to not bother to require a line length may work somewhat for Go (which tends to encourage short lines by virtue of lacking expressiveness), but it doesn't work for lots of other languages, including JavaScript.
I have big screens but side by side diffs on unformatted code on GitHub are always wrapped. Plus I often split my screen. Not everyone uses maximized windows all the time, and in a shared codebase sometimes you want to try and cater to everyone.
You don't have an infinitely long screen. You have some line length limit.
Same for tabs for indentation[1], that allows the reader to set its own indentation width depending on its own preferences.
I don't really understand why people keep using spaces + fixed line length. Old habits die hard I think.
[1] tabs for indentation, space for aligning
99% of all software projects have a de facto or de jure line length limit for a reason.
Line length limits help me to make really effective use of big screens.
Because making the line wrapping nice requires deep syntax-aware editor integration. Most languages try to be at least somewhat plain-text-editor-friendly.
Why does it need to be syntax aware?
'nice' is the key word here
I think it's just a matter of programmer habit now -- folks expect their editor to mirror the code without formatting tweaks.
It gets confusing when you realize that programmers may insert their own indentation and wrapping, and sometimes you want to make this part of the actual file, and sometimes you want to remove it before serializing.
We've actually got a bunch of automatic eslint rules, but it doesn't correct for this. Looking at the various beautifiers, I couldn't find anything that really cleaned up line spacing very well, and ended up running a bunch of sed commands to remove lines and then spacing out things manually in a "sane" way (so picky!!). Having a formatter that actually parsed the AST and rewrote the code like gofmt will be very handy if it works - I'll have to try it out.
This is a trend that I feel is going to catch on across any language that can support it. It makes trivial decisions and arguments about formatting a thing of the past.
This has been pretty standard practice in the boring ol' enterprise for at least 10 or 15 years.
Java has Checkstyle, with format-on-save support for IntelliJ, Eclipse, and NetBeans at the very least. Visual Studio supports this for all of the .NET languages. Obviously, Golang has `go fmt`. Etc.
Not to be snarky, but whenever a bold new trend seems to be really "catching on"... it's almost always a rehash of something that the enterprise was doing back in the 90's, or academic researchers were writing papers about back in the 70's. The only things that ever really change in this industry are: (1) solutions that were once impractical on old hardware become practical on newer hardware, and (2) solutions that were over-engineered in their original form come back in more user-friendly simplified forms.
Could you point me in a direction to set this up on my system?
FWIW if you use vim, this plugin is great: https://github.com/google/vim-codefmt
There were coding styles, but actually enforcing them wasn't something regularly done (linting is okay), for fear of "bondage & discipline" complaints.
Now the former "rock star" Ruby developers are embracing static typing and Ein Code, Ein Style formatting. Weird.
Personal code formatting preferences is a problem. Devs should just let go of it, once it applies automatically your preferences quickly change.
You can even wire up GitHub etc repos to automatically reject pull requests that add pointless formatting preferences someone will pointedly defend.
Why do Haskell and Elm format lists like this?
type alias Circle =
{ x : Float
, y : Float
, radius : Float
}
...and not like this: type alias Circle = {
x : Float,
y : Float,
radius : Float
}
...? The first makes me shudder with revulsion every time I see it --- that's not where commas go, dammit --- and there must be a reason, which I've never been able to figure out.The Elm style guide mentions in passing that trailing commas require a diff that adds a field to modify two lines but one, but that applies to leading commas too if you add the field at the beginning of the list.
The trailing comma is more easily missed I'd say.
https://github.com/martindemello/crosspad.elm/commit/c67deb4...
The idea is to have a common AST (with parsers for each language) which keeps track of whitespace/comments/etc and then share a ton of logic by writing writing code at the AST layer.
You can use pfff to pretty print according to your style needs. Also supports a bunch of other neat things such as search/replace.
PS: This seems to be the code for the JS beautifier in jsbeautifier: https://github.com/beautify-web/js-beautify/blob/master/js/l...
myPromise.then(() => {
// ...
}).then(() => {
// ...
}).catch(() => {
// ..
});
and if I also prefer the Lisp-like approach and hate putting: );
on its own line, am I part of the problem? What if my whole team uses that style? For example, we would write: foo(
reallyLongArg(),
omgSoManyParameters(),
IShouldRefactorThis(),
isThereSeriouslyAnotherOne()
);
as: foo(reallyLongArg(), omgSoManyParameters(), IShouldRefactorThis(),
isThereSeriouslyAnotherOne());
(100 characters wide, double indent on the continuation line. This is pretty standard for Java formatting.)Stated differently, is this meant to be very opinionated in hopes that all JS would follow a uniform style? I think that can work for Go since it was like that from Day 1. For JS, though, it has been out for so long that many teams have developed their own preferred style. I suspect that most of them would avoid an opinionated formatter that differs in a few small ways from their in-house format, if only because it seems silly to break diffs and git blame for a sweeping formatting change.
Dan Abramov said it best, when he opened an issue urging to resist the urge to add configuration: https://github.com/jlongster/prettier/issues/40
Configuration has a cost, and the product will be better and more reliable if it's limited. As it stands right now, the configuration isn't ideal for my preferences, but I'm fine with that; either I'll
* use the tool and adopt new patterns (ones which, frankly, have very little impact on anything), * keep doing this stuff manually (it's gotten me this far!) * fork the project, and bear the cost of maintaining the updated configuration myself.
The version control issues might be a blocker though. It would be neat if git had a way to do diffs at the syntax tree level.
Between this and eslint+flycheck I will feel less envious of Visual Studio Code's editor environment (if only there were a way to leverage some of its fantastic IntelliSense from emacs...)
[0] https://github.com/jlongster/prettier/tree/master/editors/em...
I think that vscode uses the typescript language server behind-the-scenes for plain javascript analysis.
So while it is another language, it is very much related.
Implementations (compilers and interpreters) should include tools/modes for converting between these two representations, and should support running programs written in the parsed format in addition to those written in the human-friendly syntax. This isn't asking much: the difficult part is parsing the human-friendly syntax, which implementations must already do.
The benefit is that we don't end up with a mismatch between what the compilers/interpreters accept, and what the other tooling accepts (linters, formatters, doc generators, static analysers, search engines, syntax highlighters, use-finders, go-to-definition, refactoring tools, etc.). This would, incidentally, allow people to read and write code as s-expressions, but that's not the point; it's just about representing concrete syntax as closely as possible (plus arbitrary annotations, e.g. for line numbers, etc.), whilst exposing the structure in a machine-friendly way.
This would also make it much easier to make new tools, extend existing ones, and improve practices, e.g. like syntax-aware diffing, standard code formatting (i.e. tools like this one), version-control-friendly representations, tree-based editors, more powerful navigation in editors and IDEs, etc.
https://github.com/esbenp/prettier-vscode
It is my first VS code extension so have no idea if it is correct way. Feedback is welcome.
EDIT: Looks like they may be open to it though. https://github.com/jlongster/prettier/issues/12
“You can use standard --fix to automatically fix most issues automatically.
standard --fix is built into standard (since v8.0.0) for maximum convenience. Lots of problems are fixable, but some errors, like forgetting to handle the error in node-style callbacks, must be fixed manually.”
I'd disagree. See https://www.youtube.com/watch?v=wf-BqAjZb8M
For tabs/spaces and semicolons I'm considering if we should support those options. We need to figure out the goal of the project: is it to converge on generally a single format, or is it to provide formatting options for a few large groups of people that have different opinions. There are issues on the project discussing this right now.
Perhaps stick to just one, as seems to be the case right now (whichever it is), and mention in the documentation that those wanting the other option are free to e.g. write a patch/fork on GitHub/postprocessor/etc., with the understanding that such support will eventually get merged in iff it's actively maintained for some amount of time, its author/maintainer is active in general development and maintenance for the project, and there's significant community adoption of such a patch.
Whilst not perfect, this sort of approach might better determine who actually cares enough about this to offset the community-splitting effects of allowing both; compared to a general accumulation of online grumbling.
As was once told to me, my code should look like your code and yours like mine.
Honestly, why does this matter at all?
function makeComponent() {
return /*test*/ { a: 1
};
}
into function makeComponent() {
return /*test*/
{ a: 1 };
}
From my own experience writing a formatter, comments are very difficult (if one wants to preserve them). I guess the author should reparse the generated code to ensure the AST hasn't changed.There are a very few subtle bugs like the one you mentioned above. As noted in the readme, this is a beta. It's certainly more than an attempt.
function doSomething() : int {
return myPromise
.then(() => {
// ...do thing
/* what */
})
.catch(() => {
// ..do other thing
/* the heck */
});
}
becomes function doSomething(): int {
return myPromise
.then(() => {})
.catch(() => {});
}
Still a very nice pretty printer, though. JSX is<strong>supported</strong>
Note the lack of space after the word "is".This will add semicolons all over your code depending on some rules, for example, if you write a return, and write the value you want to return in the next line, it adds a semicolon after the return AND after the value, effectively ignoring the value in the end.
return
myValue;
becomes return;
myValue;
It's a bit meh, because you have to remember these rules even if you use semicolons.This would not catch changes like
--val x = [1]
++var x = [
++ 1,
++];
as expected, while it would give you a slightly more verbose diff if the change was --val x = [short(), short()];
++val x = [
++ a_really_long_name_that_pushes_the_line_length(),
++ a_really_long_name_that_pushes_the_line_length(),
++];huh? I've been using emacs with WebMode + tide + FlyCheck and it supports jsx just fine. Moreover, tide[1] provides great support for plain javascript (I've used it extensively on a ES6 codebase with great results).
Would be cool to have this as a little website to try out, too.
It would be nice if the JSX in React's render could be formatted nicely because a lot of times the code get's so long that formatting and indenting becomes inaccurate.
I like to have multiple windows on screen at once, and will often adjust them to make best use of screen real estate. I tend let it soft-wrap the text, but I'd love it if it would do soft-wrapping... prettier.
I got a quick feel for it with this in a terminal:
watch -n 0.1 'prettier --print-width $COLUMNS index.js'
... and unfortunately didn't really like it. Maybe in a pinch, but it still ends up feeling really cramped trying to read column in low-width terminals. At the other extreme (having the terminal be the only window on a 21:9 ultrawide), some of the lines get so long that you'd want a max anyways.e.g. :
function something()
{
}
i.e foo({ num: 3 },
Honestly, we may never do it. But we're still figuring out if the ultimate goal is to keep to a single style or not.
foo(function() { return {num: 3}; });
I used that a lot when writing code that used D3. But since arrow functions, I hardly ever use a single line block statements, so this style is obsolete.gofmt will fix indentation, spacing, and so on, but it will generally preserve structure. For example, this:
a:=Foo{value:42}
becomes, of course: a := Foo{value: 42}
But! This: a:=Foo{
value:42,
}
becomes: a := Foo{
value: 42,
}
That's because gofmt can't really pretend to know that it knows better than the developer here. Sometimes code does need to be loose (like in a DSL or a big declaration, or a test (which must be readable), or similar). Sometimes it should be compact.This means you never have to fight gofmt. Never once have I disagreed with its decisions.
gofmt works this way because its ruleset isn't exhaustive; it says that, yes, an indent must happen after a hanging "{", but the ruleset doesn't say that a line break must happen. If there's a line break, let's stick with it.
Prettier seems to have a strict normalization approach: AST goes in, canonical form goes out. For example, I tried this fictional piece of code:
performAction({
type: "thing",
value: 42,
owner: user,
bucket: root,
path: p
});
It's turned into a much less readable blob because it happens to fit on a single line: performAction({type: "thing", value: 42, owner: user, bucket: root, path: p});
Unfortunately, in every instance where I indent the code in this manner, it's for a specific reason. I wanted it on separate lines; the fact that it happens to fit on a single line doesn't matter at all. Prettier overruled my carefully indented code.I often format code in a specific way for regularity: Every chunk in a block should have the same format, because each chunk is an instance of the same pattern. For example, I might have something like:
const RULES = [
{
type: "boost",
fn: node => increaseImportance(node),
weight: 0.5,
dependencies: ["priority", "rotation"]
},
{
type: "eliminate",
fn: node => nil
},
// ... lots of rules ...
]
Prettier's output generates inconsistency: const RULES = [
{
type: "boost",
fn: node => increaseImportance(node),
weight: 0.5,
dependencies: ["priority", "rotation"]
},
{type: "eliminate", fn: node => nil}
];
To quote the immortal Trump: No way!Maybe one could use some advanced heuristics to find an optimal balance between width vs. indentation vs. compactness; for example, in the above array, a clever formatter could see that it's an array of object literals, which means that it should prioritize regularity over compactness. If it's an array of something simple (like numbers, but not numbers with trailing comments), it can go compact. Maybe.
I don't use Prettier at the moment, but I know the strictness would drive me nuts. I predict that Prettier is going to cause a lot of frustration and heated discussion as a result of the one-size-fits-all approach. I don't think a canonical form for everything even makes sense; people need regularity (and no surprises), but not at the cost of readability.
I am trying to write my own formatter at the moment, and this is exactly the property I want, also. I'm working in Haskell, and one example is case statements. There are two ways I'd like to format them:
case foo of A -> ... B -> ...
or
case foo of A -> ...
B ->
...
The first is when every case fits on one line, but the second is when at least one case needs to span multiple lines. I don't want this:case foo of A -> ...
B -> ...
I haven't yet come up with a good way to do that heuristically and efficiently, but maybe just looking at how the original code is formatted would be enough! Thanks for providing some food for thought :)