Delimiter-first code
arogozhnikov.github.io
arogozhnikov.github.io
Indeed, feels super unusual and tooling is lacking, but behold the concise rationality:
; const url
= "..."
; const contact
= { "firstName"
: "John"
, "lastName"
: "Smith"
, "address"
: { "streetAddress"
: "21 2nd Street"
, "city"
: "New York"
, "state"
: "NY"
, "postalCode"
: 10021
}
, "phoneNumbers"
: [ "212 555-1234"
, "646 555-4567"
]
}
; const result
= sendData
( url
, contact
)
; rejoice
( result
)
Beauty. N.B. those opening and closing delimiters "connected" by breadcrumbs of separators. Isn't it lovely?Most developers I know will immediately lit torches and grab pitchforks if you show them this.
(Personally, I'm using it sometimes for private doodling, and while reformatting obfuscated or exploring some complicated codes; mostly just to improve own reading comprehension.)
If you relax the rule to allow "single simple value" to be in the same line (so break the rule for "=" / ":" on the start of the line), it could look lie this [1]
The reasoning behind that is that it produced the cleanest diffs with source control - especially if you consider changing the first/last/only entry in a collection.
Consider the diff from removing the first phone number from your collection above, and compare that to removing it from
"phoneNumbers": [
"212 555-1234",
"646 555-4567",
],By which metric is that rational? IMO it is utterly unreadable.
> Most developers I know will immediately lit torches and grab pitchforks if you show them this.
Yes I will do just that now lol
Prettier was invented so we don't have to deal with this bullsh*t anymore. What a gift of god
In contrast of the conventional ("irrational") zig-zag pattern where: - starting delimiter is somewhere in the middle of the line (if not in Allman and similar style); - separators are mostly at the right end of the line, sometimes pretty far, depending on line length; - closing delimiter is mostly on the left. - Plus separators and delimiters can be scattered among single line, if they "fit" there.
Again, I'm not telling this is "good and YoU ShoULD USe thAT!1!", only that compared to rules of other conventions I see least amount of "rules" and "exceptions" in this. And it is usable only in syntaxes with separators (JS to some degree, JSON, CSS), not for languages where ";" must finish even the last statement in a block (PHP) - there you'd have dangling two characters.
As for readability, I concur, yet I'd be very cautious of raising judgments like "it is utterly unreadable". I know that readability is largely matter of habit and I'd not be surprised that if we lived in such "Haskely" alternative universe where this convention was a norm from the start, we'd probably scream in terror when confronted with "Prettier" ("Stroustrupy") conventions.
That "emphasis" you are talking about is probably just that they seem unusual so they draw attention.
I'd change careers if I had to format or read code like that...
The a) b) c) and 1) 2) 3) argument seems just confused - those are not delimiters.
The author claims that we don't need a terminator - but what do you do in case of two subsequent sequences ?
Author says that YAML is "delimiter-first", yet there is no \n at the beginning of a YAML file.
In the end it all comes down to putting an extra comma in front of everything you write at which point you might as well just parse everything backwards to achieve the same result.
Once I take out all the circular logic out of the theory I simply fail to see any benefits of this approach.
I've been experimenting with how this type of syntax could be used to transcribe Lisps, following on from the like of 'Wisp'. The results are fairly interesting, in my opinion, though I've mostly been fiddling with the feeling of writing and reading so far, rather than actually settling on a syntactic combination that's coherent. I'm using bullets rather than numbered lists, since sequence is basically always implied in code context.
https://gist.github.com/armstnp/bb2a88bcb053d2195f42c60a0cf1...
https://gist.github.com/armstnp/3fab1a0933c0da273525f77799de...
Hmm... I don't know about that. They fit many of the definitions I can find online:
> A delimiter is a sequence of one or more characters for specifying the boundary between separate, independent regions in plain text, mathematical expressions or other data streams. -- https://en.wikipedia.org/wiki/Delimiter
> a character that marks the beginning or end of a unit of data -- https://www.merriam-webster.com/dictionary/delimiter
Edit: Have to revise that, it made me rethink something: maybe I should throw endOfLine as a token out of my parser generator engine, and instead introduce a startOfLine token.
Of course, that means that sometimes you have to look at code in the latter while you prefer the former. For instance, when you examine a merge conflict. But for the most part, in your everyday programming work, you only see code formatted the way you like.
Of course, this assumes that it's possible to programmatically translate between arbitrary different code styles. This is probably true when we're only talking about differences in whitespace, but the story become quite different when we also consider things like naming conventions (different case styles (underscore vs. CamelCase, first letter small vs. first letter big, etc.) and other things.
Not sure if it's still the case, but the Linux kernel used to have a script checkpatch.pl that was a few thousand lines of Perl, dense with regex, to check that patches conform to kernel style guidelines. And those guidelines are basically just K&R C style.
When I was a teenager I really wanted to get into contributing to the kernel, so I did this babby's first patch, fixing the code style for this rather large, extremely messy 3rd party driver in the tree. Must have taken me a couple days; hardly any single line conformed to the guidelines. Sent the patch in and Greg KH replied saying thanks and all, but we were planning on just deleting this driver anyway, so no merge.
Somewhat embarassing intro to open source contribution :D
The benefit of this is that it almost cannot go wrong (unlike older formatters which might stumble over a line of code and accidentally write invalid or incorrect code). The disadvantage is that there are usually fewer configuration options, so if you don't like the formatting, you might just be stuck with it anyway.
Instead of arguing about code format styles, why not change our languages (or at the very least development environments) so that a list (for example) gets parsed, analyzed, and turned into an internal object, and doesn't remain merely a "dead" string of characters.
Once lists are "objectified", we should be using tools to modify those objects (similar to how you'd add/remove rows from a spreadsheet, or edit the DOM in your browser's dev tools). How those objects get presented visually to the programmer as source code is irrelevant at that point: Want to show the list a a single line of items concatenated with commas? No problem. Want to see each item on its own line, without commas? Sure. The visual representation of lists just become syntactic sugar, while the actual list itself is embedded in the development environment as an internal object, much as your HTML pages become parsed by the browser into the DOM which you can inspect and modify via the browser's dev tools.
We should be coding to build "smart objects," not writing strings of text to be endlessly curated.
The way it should work is that when loading a file,
your editor formats the code the way you feel most
comfortable working with it.
This doesn't seem practical if you have to reconcile external sources of information with the file - for instance, if you're looking at line numbers in a stack trace from production. -- no commas
SELECT employee_name
company_name
salary
state_code
city
FROM employees
Or: (select (employee_name
company_name
salary
state_code
city)
from employees)I don't think your example looks very s-expressiony with the FROM keyword in the middle of the sequence. To me, s-expr means the keyword in the first slot labels the form to interpret the rest of the sequence. The outer form should either have positional slots to implicitly interpret each part by different rules or an unordered set of slots that will each have their own labeled form. Consider instead something like:
;; (QUERY select from where)
(query (e1 e2 ...) (r1 r2 ...) (e1 e2 ...))
;; (QUERY parts...)
(query (select e1 e2 ...) (from r1 r2 ...) (where e1 e2 ...) ...)
Also, existing SQL syntaxes have some places where parenthesis has a special syntactic purpose and others where it is optional. It is rather convoluted to define unambiguous parses for some of these. For someone familiar with any of this, it would be hard not to bring in false assumptions about how a hybrid language would work that brings in parentheses for other structuring purposes.
You don't care much about syntax, do you?
It makes for cleaner edits and diffs.
If you are designing a syntax that does not rely on line breaks like Python, reach for "introducers" instead of "delimiters". Commas make awkward introducers even though people can and do get used to it.
result = a==1
|| b==2
|| c==3
;
Vs result = a==1 ||
b==2 ||
c==3 ;
and result = data
.getA()
.getB()
.getC()
;
Vs result = data.
getA().
getB().
getC();
Personally I prefer the first examples, because for me it's easier to "group" the lines. foo = here_is_some_logical_condition ||
and_another ||
(this_term && that_term) ||
exception_case;
vs foo = here_is_some_logical_condition
|| and_another
|| (this_term && that_term)
|| exception_case;Records: https://github.com/janestreet/base/blob/master/src/avltree.m...
If statements: https://github.com/janestreet/base/blob/master/src/string.ml...
Lists: https://github.com/janestreet/base/blob/master/src/string.ml...
Function calls: https://github.com/janestreet/base/blob/master/src/string.ml...
Type definitions: https://github.com/janestreet/base/blob/master/src/array.ml#...
It's not always used, shorter snippets tend to be all on one line.
[ 1 2 3 ]
In particular syntaxes that require a delimiter and forbid a trailing delimiter require much more verbosity in code-generators.This is maybe the only place where bash differs from other syntaxes but (IMO) clearly got it right.
I'm also reminded that Clojure uses a compromise where there are no delimiters, but a comma is treated as whitespace.
[edit]
Also this comes up from time to time in the Nix language; I was complaining with someone and we both were annoyed by the inconsistency; lists use no delimiter while sets (i.e. dicts) use a semicolon. I thought that sets should have no delimiters, he thought that lists should, but we both agreed that the inconsistency was jarring.+
Of course, if sequences are callables that reproduce the original sequence with the argument appended, then these kind of merge, but with empty list as a prefix instead of list bracketing.
Few complaining about significant whitespace think:
foo bar = 3
Should assign 3 to a variable named "foo bar" because otherwise it's significant whitespace. The lexer already splits tokens on unquoted whitespace for most languages, so requiring a comma in lists solves ambiguity.The only real argument against it would be that this looks confusing (though it's still unambiguous!)
[ 1 2 + 3 4]
so one would either have to live with that or parenthesize multi-token expressions in a list. Note that due to Nix's weird function syntax, parenthesizing function-calls in lists is already required there, but c-like languages don't have this problem.
People who hate on “significant whitespace” tend to have a fairly arbitrary set of conditions for which whitespace being significant, and in what way, bothers them.
Argument 1 is not about whether tools can handle the comma-first format well or not (which is what the counter-argument focuses on).
Instead, the argument that tools can themselves solve the problem that comma-first format is used to work around (so there's no need for it).
That is, the "anti-comma-first" argument is not that "we shouldn't use comma first because tools can't easily handle it".
It's "the thing people attempt to solve with comma-first, namely not accidentally forgetting a comma, tools like linters can warn us about, so it's not something we should change our natural formatting to manually solve".
If you instead compare:
select
foo,
bar,
fred
and: select
foo
,bar
,fred
it's arbitrary, and a shame there's no trailing (resp. leading) commas allowed in either case.If the marker is "-", EBNF might look like this:
leading-marker-list = *("-", element)
Comma separated would look like this: comma-separated = [ element, *(",", element) ]
This is not meant as an argument for or against; just an observation.Edit: Amended comma separated EBNF to account for an empty list.
Hence it would be:
leading-marker-list = [ element, *("-", element) ] while (test("-")) { element() }
vs. do { element(); } while (test("-"))I have written recursive descent parsers in the past, and if I remember correctly, the way I set things up, I had to handle leading list elements specially. It's possible that this was unnecessary and I was simply blinded to a simpler way of doing things because I was following the EBNF too strictly.
Edit: I noticed that we're not actually accounting for an empty list here. That makes things slightly more complex:
comma-separated = [ element, *(",", element) ];
Your code would need to test if there is an element or a list terminator before processing the element. sections = [
"one"
"two"
"three"
]
and separate with commas when you want to write something on the same line sections = [
"one"
"two", "three"
] sections: [
"one"
"two"
"three"
]
and sections: [
"one"
["two" "three"]
]
exists already and is named rebol (and their derivatives) irb(main):001:0> sections = %w(one two three)
=> ["one", "two", "three"] SELECT , employee ,
, company ,
, etc ,
Or maybe even: SELECT , employee ,
,
, company ,
,
, etc ,
The first comma asks the data to confirm comma-separatedness, the second acknowledges the confirmation, and the third is the actual payload for the comma separator.(And let's just go ahead and disallow trailing three-way-handshake commas for fun. :)
We have a special word for end-of-item token: terminator, but no startinator or any similar word. I see some irony in this.
Maybe "introducer"?
Delimiter-last is just more intuitive, it’s how we do it in natural language (I like bacon, butter, eggs, and cheese).
There are some benefits to delimiter-first, but its just not worth the effort to switch.
for (x in items; let comma = false) {
if (want_to_print(x)) {
if (comma)
print ", "
print x
comma = true
}
}
A delimiter-separated list should be regarded as all but the first item being preceded by the delimiter, rather than all but the last item followed by a delimiter.If you that the latter view, you write silly code which tries to guess whether more items are going to be printed in subsequent iterations, and thus whether to generate a trailing comma now.
The delimiter goes directly after the items because it means ‘but wait, there’s more’, and so mentally puts you into a mindset of being prepared for the next item. If list-like delimiters are placed at the start of lines, the reader is forever in a heightened state of ‘is this thing over or not’ as they go on to the next line.
I’m not arguing that this is world-ending or anything, but I do think it’s what makes me feel ‘uneasy’ when reading the examples.
Do you seriously propose writing something like
2,7Sdata2,3Cfoo2,3Cbar2,4C12343,11Chello world2,7Cgoodbye2,7Chunter23,26Cswordfish is a red herring
instead of data: string[7] = {
"foo",
"bar",
"1234",
"hello world",
"goodbye",
"hunter2",
"swordfish is a red herring",
}
? Thanks but no thanks, almost nobody likes to write Bencode by hand, just trust me.Edit: expressions vs statements.
SELECT
employee_name
, company_name
, salary
, state_code
, city
FROM 'employees';