But if the syntax was separate from the underlying representation, couldn't you just have your editor open it the way you want?
But if you just always indent by n tabs, and align with spaces for visual alignment, then it's not a problem. This is what e.g. gofmt does.
(This isn't the best example of what I'm getting at, but it's easy to grasp. A stronger example would be the use of ternary operators with multiple lines of JSX. There are many "micro-syntaxes" or "micro-layouts" that help readability and form common patterns in code bases which don't play well with tabs-for-indents, so when you use tabs, you have to train your team to develop and accept novel versions of these "micro-layouts".)
In other words, instead of:
def my_func(self,
param):
You always do something like: def my_func(
self,
param,
):However, regarding "mixing two different invisible characters in the same file" - you can, and should, configure your text editor to display tabs in a visually distinct way from spaces. Kate (my favorite) shows them as chevrons. Geany and Notepad++ show them as arrows. Not only does this make it easier to scan vertically along an indentation level, but doing so also makes it glaringly obvious if a pesky space has snuck in, resulting in a jarring misalignment.
While I agree with this in general, I find making the code nicer to read to be more important that avoiding diff pollution. Especially when that diff pollution can be avoided with a simple "ignore whitespace" in most cases.
That being said, there are other cases where I want to align things and "a specific size indent" doesn't work
SQL
SELECT this_column,
that_column
FROM tablename
WHERE this_column = 1
AND that_column IS NOT NULL
Comments /*
* Blah blah blah
* Note: this is interesting because (enough text to want newline)
* and also that blah blah
*/
There's lots of places where "alignment" makes sense. SELECT
this_column,
that_column
FROM tablename
WHERE
this_column = 1
AND that_column IS NOT NULL
(The AND token could go either where it is, or at the end of the previous line, depending on your preferences - I prefer the clarity of putting the operator first here over the alignment of the column names, but I can see both sides.)It might not be your preference, but that's always going to be the case with formatting.
With the comments, that's a good point. I'd generally prefer to use double comments for anything that spans multiple lines, but if you're creating a formatter you need to handle any syntax thrown your way. You could try formatting it without the leading space, but that would be very unusual, or keep the leading space and have it as the one exception to the rule. Or just not format comments, I guess.
Code is so incredibly hard to mentally grasp and every mental overload should be omitted to reflect on the logic.
There is a reason why there is one code basis and this should always be curated by a linter to uniformly enforce a standard.
There is still plenty of room for style and code organization.
I witnessed first hand many trench wars around seemingly small things like curly brackets in IF statements dealing with the question of one white space or none, because it appealed to personal preferences and before ESlint people would go to great length reformatting hundreds of LoCs just to get their right feeling of code syntax.
Weird. And git -diff was massive, as well as the code reviews.
(Un)Happy times. :D
You can configure ESLint to use prettier, but in my experience it's usually easier to just run ESLint separately to prettier.
>Maybe we'd have editors and viewers where you could configure the syntax however you wanted
I took this to mean that, in this fantasy universe, you could make any source file look however you want. Like tabs vs spaces and pure html vs html-with-css, this is about separating meaning from presentation. Is there a good reason to force the same visual representation on everyone?
Leaky abstractions. If there are two different visual representations then sometimes someone using one representation will need to think about the other.
* Would such bytecode/binary be considered "source code"? If no, then bye bye Open Source Definition (OSD). Distributing such thing would not be considered open source.
* If it is considered "source code", then is it considered "obfuscated" source code? If yes, then bye bye OSD. Distributing such thing would not be considered open source.
* If it is not considered "obfuscated" source code, then what's the difference compared to, say, JVM bytecode? Is it only that an IDE exists that can modify such bytecode in a fancy way?
* In that case, if such an IDE is created that modifies JVM bytecode directly, and a program is created using that IDE modifying the JVM bytecode[1] directly, and such "source code" is distributed, the would that bytecode be open source? Since, after all, it's "the preferred form in which a programmer would modify the program" (and therefore not obfuscated).
[1]: Or maybe some superset that is mostly JVM bytecode but it has some opcodes for higher level niceties.
I’m all for keeping consistent flat utf-8 files. I’d hate for my code ultra simpleminded, possible to pen test with a pen and paper, python and sql code to be wrapped in a god awful xml or json or proprietary db markup and object hierarchical model of what code should look like.
Like I imagine trying to check in a jupyter notebook but worse.
For instance tabs vs spaces was decided and text editors accommodated this and despite what may be someone’s personal preference a uniform decision was made.
Humans can learn. We should use accessible formats and push for standards and keep those readable.
It would definitely be the norm that there were languages that were dark or light mode, and both sides would be convinced they were right.
Do you group things like:
do_stuff()
more_stuff()
other_stuff()
moar_function()
Or: do_stuff()
more_stuff()
other_stuff()
moar_function()
Or: do_stuff()
more_stuff()
other_stuff()
moar_function()
Grouping code by a single blank line can make a big difference. This is a bit of a silly example, but there's tons of not-so-silly examples. You can't really represent that in an AST.Line length is another. infinite line length doesn't work as screens aren't infinitely long. Automatic wrapping doesn't work because you want to break at specific points. An example is something like:
window = XCreateWindow(display, XRootWindow(display, screen),
center_x, center_y, size, size,
4, // border width
vinfo.depth,
CopyFromParent, // class
vinfo.visual,
CWColormap | CWBorderPixel | CWBackPixel | CWOverrideRedirect,
&window_attr);
Cramming as much as possible on as few lines as possible is just not going to work well.There's tons of cases.
That's why no one does it. Because it just won't work. Everyone will hate it.
That said, if you use classes, descriptive names, and appropriate comments, it won't matter how you group because your code will be self explanatory.
Finally, with today's wide-screen monitors on desktop, line length is less of a worry. Problem only arises when reading code on mobile devices.
The last workplace I was at had a soft wrap around 80 characters, but we upped that to 100 when functions and methods became almost vertical.
window = XCreateWindow(display, XRootWindow(display, screen), center_x, center_y, size, size,
4, vinfo.depth, CopyFromParent, vinfo.visual, CWColormap|CWBorderPixel|CWBackPixel|
CWOverrideRedirect, &window_attr);
And that's just not better. It's much much worse. Only a human can judge where it makes sense to break things (or not break things).Completely forbidding "paragraphs" would be atrocious. Almost everyone will hate it. Just like everyone hates walls of text in comments.
And in the end a huge amount of complexity is added throughout the entire stack (from editors to grep to sed to code hosting) for very marginal gain with significant trade-offs.
Also, why does your comment have a paragraph break after each sentence except after the third? Because of semantics and the related readability. The same is true for source code. Readable layout does not depend on syntactical structure alone.
did you buy a large monitor just to full-screen a single text file?
1. Breaking at specific points is something that can be specified by the pretty-printer of the _viewer_ you are using. Think of existing auto-formatters and imagine that they're working over the view instead of the persisted form.
2. The AST can have pointers into advisory data (or the other way around, if desired, the "program data" can include the AST, but also other things as well) to note that there is an anonymous region here (C# and friends already have conventions for this _for the source code_ that Visual Studios understands - look at `#region` comments). This would let viewers choose their preferred representation for a region.
And including tons of extra info in the AST means you've just got the same as a text file, but in a more awkward and obfuscated format.
Also what you want can already be done right now. Converting your source code to AST and then formatting it how you want is already easy to do. But your coworkers will hate you, because you will lose all these small details that are actually quite significant. And there is no way to get them back once they're lost, so your coworkers all using some formatting tool is not going to help.
<visual-group>…</visual-group>
Also, I wish that at least in current editors there would be a way to render \n\n as a half-height line. Full-height empty lines are too bold.
Automatic wrapping doesn't work because you want to break at specific points
You really want to have a set of buttons that switch between:
- a line of arguments
- a block of arguments
- a line/block of only non-default valued arguments
- arguments sorted by name
- …
And a hint on a default representation, with some default heuristics.How does it know that x, y, width, and height are semantically grouped and best put on the same line? It doesn't. (from the other comment)
Look higher, parameters can be grouped at the declaration level. <params><related-params name=“coords”>…</>…</>. Now you can render a nice frame around these in block mode. Or not, depending on local renderer settings.
Everyone will hate it.
There’s always a way to make everyone hate something, especially if the solution is clueless about its problem. It doesn’t mean it should be done this way. Experiment and evolution could make it work, we just have to let people try instead of dismissing it so confidently.
I don't see why it couldn't be done though, I think it just hasn't been a priority. Heck, you could have 100 different users collaborating in 100 different "languages", and so long as they serialized to the same AST and back, none of them would ever have to see the atrocious syntax which the other users prefer. Their editors and browsers could just render everything according to their users' preferences.
Edit: it appears that Unison has an issue for this feature: https://github.com/unisonweb/unison/issues/499
You could then decompile to some alternative syntax, but you'd lose any idiosyncratic formatting represented by the compressed diff.
Last I checked, Java AoT compilation precluded runtime re-optimization, though I presume they've fixed that by now.
Last I checked, they both used stack-based bytecode, which typically takes longer to JIT and results in slower native code than a compressed SSA / control flow graph (see the SafeTSA papers).
Though SSA is deferred to JIT/ILC instead. In either case you get the access to all the actual low-level bits when you need to. No other portable target lets you do that.
foo
.bar()
now you need to either add syntax to specify that you’re “continuing” (python), play tricks for the parser (Go), or have unreliable magic biting you in the ass half the time (javascript). foo.
bar()
And that will work fine. At least in Go.You can endlessly argue what location is better for the ".", but it really doesn't really matter. Regardless, you don't really need to do parser tricks.
It also looks like shit, because now you have to check the end of the previous line to know whether it's a method call or a function call which was indented in incorrectly.
I do see what you mean, of course, and am aware of the arguments but I've never found lack of non-whitespace expression terminator even remotely confusing. I can clearly see indentation in my periphery signalling that it's multi-line. I can also say that in ~13 of using exclusively using semicolon-less languages, I've never had or even heard of a problem caused by someone misunderstanding where an expression ends.
Anyway, this is getting pretty bikesheddy, lol but ya, to each their own, obviously.
Plus I don't think I've ever seen a case where semicolons made the code harder to read (or write).
Note that I said 'statements', not 'expressions'.
A lot of the confusion here (and maybe yours, too) stems from this difference. In Rust, (almost) everything is an expression by default, and you turn it into a statement by adding a semicolon. This allows you (and the type checker) to very neatly distinguish between expressions and statements, which is great. It's a very nice and elegant approach imo.
Method chaining is also common in Rust, because builders are common, and chains of iterator adapters are common, and chains on monadic structures (option/result) are common, … having every line break implicitly insert a `;` would be horrid.
fn five() -> i32 {
5
}
and this isn't? fn five() -> i32 {
5;
}
I can't tell if I'm amazed or terrified. fn five() -> i32 {
return 5;
}
fn five() -> i32 {
5;
return ();
}
semicolon changes the expression from returning the result of the expression to running the expression and returning the unit type. if you accidentally do that and specified a non-unit-return-type in the function signature, the type checker tells you about it: error[E0308]: mismatched types
--> src/main.rs:1:14
|
1 | fn test() -> i32 {
| ---- ^^^ expected `i32`, found `()`
| |
| implicitly returns `()` as its body has no tail or `return` expression
2 | 5;
| - help: remove this semicolon to return this value
which is also pretty clear about the solution> I can't tell if I'm amazed or terrified.
The rule is very simple and obvious, and the compiler will yell at you if you get it wrong.
It's also very useful and even critical of how expression-oriented the language is: an `if/else` or a match statement must typecheck, all branches have to have the same type. It's obvious if you're using it as an expression, but it doesn't go away if you're using it for the side-effect (as an imperative conditional/switch) and then things can get more dicey as the expressions in each branch can have different types. `;` solves that by making every branch `()`-valued.
fn five() -> i32 {
let a = if (true) { 3 } else { 6 };
return a + 2;
}
Pure fun!And though for abs while don’t (they always return ()) `loop` itself does, you can `break` with a value and that’s what comes out the loop.
let x = {
3
};
let y = {
3;
};
assert_eq!(x, 3);
assert_eq!(y, ());
i think it'd also mean having to parse whitespace or newlines without something like that?I'm playing around with this in the compiler explorer and I'm now even more confused why you want this.
let x = {
println!("Don't mind me in the struct definition");
3
}
assert_eq!(x, 3); let dx = {
let prev_x = x;
x = get_x();
x - prev_x
};
often, it's slightly better cpp style scoping blocks if nothing else? there are tons of other little QoL things it enables though, but they're all going to be little ergonomics things that only seem worth it if you've used the language for awhile spawn({
let a = a.clone();
let b = &b;
move || {
// do something with the a you cloned and the b you borrowed
}
})It's not so much that this ability is specifically put in for some reason. It's just something that falls out of several other things.
Rust chose a "curly braces and semicolons" syntax because that's the sort of syntax that is normal in the sorts of PL spaces Rust wants to be used in. I am not sure exactly why being expression-oriented was chosen, but if I had to guess, it would be due to that just generally being considered a better option among many people at the time it was chosen.
So okay, you want expressions, and you want semicolons. Therefore, "you separate expressions with semicolons" is a pretty natural outcome. And since we're expression oriented, "a block is an expression that evaluates to the final value" is near tautological. And since it's an expression, it can go anywhere an expression can go.
Not being able to do this would mean creating specific restrictions against it, and then having to memorize when things don't follow the usual rules. That's more complicated than just letting expressions be expressions.
(also your let is missing a semicolon)
It works with RAII to create a temporally isolated resource scope, think context managers / using / try-with-resource; it provides a scratch space where you limit collisions with the rest of the function and avoid the risk of mis-reuse of those values; and before non-lexical lifetimes it was critical to limit borrow durations. It's also routinely used in combination with closures, to prepare their capture. Similar to C++ capture clauses, but without special syntax.
It's essentially a micro-namespace inside the function.
You can use this to e.g. acquire a Mutex guard, move/clone something out of the mutex, and ensure it gets dropped as quickly as possible.
let x = {
let items = vec![1, 2, 3];
items[1] // copied out of the vec since usize implements the Copy trait
// compiler inserts drop(items) here
};
// items is no longer valid
assert_eq!(2, x);
It also comes up in if/else blocks, which have exactly the same syntax and semantics (i.e. they are expressions, not statements): let condition = true;
let x = if condition {
println!("condition is true");
5
} else {
println!("condition is false");
10
};
assert_eq!(5, x);
edit: and of course, function blocks work exactly the same way! It's neat.Personally, I much prefer the design Go uses (where semicolons are implicitly added at the end of newlines following an identifier, numeric or string literal, keyword, or operator).
If you follow those rules, you will see that this:
if true;
{
fmt.Println()
}
else
{
fmt.Println()
}
Will get rewritten to: if true;
{
fmt.Println();
};
else;
{
fmt.Println();
};
And that won't work. That's why the braces need to be as "} else {". and "if .. {".It used to be that the compiler gave some pretty confusing errors about semi-colons on this, but it seems that's been improved now.
JavaScript has similar semi-colon insertion by the way, but with some different (more confusing) rules.
What if the AST is persisted as S-expressions, but then you have a different syntax to edit it? Algol-ish, Pascal-ish, C-ish, Python-ish: choose your poison (or even support multiple poisons and let the developer pick the one they prefer?)
This was actually the original plan with Lisp. Lisp was originally supposed to have two syntaxes, S-expressions and M-expressions, with M-expressions being Algol-like. However, the implementation of M-expressions was delayed, and people got so used to using S-expressions directly, they decided M-expressions were unnecessary and they were never implemented in mainstream Lisp. They were implemented in the Lisp 2 project, but that ended up being an evolutionary dead-end; various attempts at the idea have happened since but none of them really took off.
Lisp purists will argue M-expressions are unnecessary and S-expressions are all you need. However, S-expressions can make the language more foreboding to complete beginners, and even among experienced programmers, a decent percentage find them seriously off-putting. Maybe if the M-expression idea had been pursued more seriously, Lisp might be more popular today.
I would argue that S-Expressions turned out to be extremely practical for developing Lisp software. A layer with different syntax makes it more complex to use.
> Maybe if the M-expression idea had been pursued more seriously, Lisp might be more popular today.
Maybe, maybe not. People have discussed it endlessly, but there is no conclusion.
Currently some people in the Racket community try to make Racket more "popular" (widely used, ...) by providing a new syntax.
When you start thinking of your program as a nested data structure instead of a text document, you start thinking more clearly - you can manipulate this tree representation directly. This makes molding your program significantly faster, closer to the speed of thought. Wholesale refactoring a function might only take a dozen keystrokes to move the forms (not the chars) around.
I have a hard time going back to the syntactic busywork of manually placing ascii characters like a caveman.
But lisp did not aim for that from the start, as the GP hinted McCarthy intended S-expressions as a data description form, m-expressions were supposed to be the executable language.
Many editors are already altering what is stored on disc before presenting it to you - type annotations, code folding, git info. I think we could do a lot if our default storage was the semantic representation of the code.
Imagine writing a blog about your favorite CSS features. I mean, if you use Word at all for writing, wouldn't that come handy?
But Word actually is not the primary goal here. grep or awk may be more important - I've created ad-hoc tools for ad-hoc quick-and-dirty tasks on code a number of times, and the fact that code is just text helped a lot. Taking that away would make me spend time on figuring out how to hook up text tools to that format.
> Many editors are already altering what is stored on disc before presenting it to you
True, but the difference is between being able to enrich commonly understood format, and not having a common format between different tools at all.
I do feel sad that we're still restricted to tools like grep and awk that treat everything as dumb text, and I think it holds us back hugely. I much prefer the powershell model of everything is an object. You can easily extend it with C#.
I appreciate that the unix philosophy has gotten us to incredible places, but I do wish we could rip off some of the training wheels and take advantage of the computing power that we have available to us that was just unthinkable. My mobile phone is more powerful than could possibly have been imagined when we make restricted itself to tab as a delimiter, and I'm currently using it as a coaster to save my table from water marks.
And you will have easier time getting your changes merged into dotnet/runtime than into Python.
Even simply "sharing" within your own systems, like copying blocks into notes or another program, would be a lot harder. Maybe I'm not knowledgeable enough here and this wouldn't be as thorny as it seems.
Your AST is what EMF calls a "model". By default the "backend" and ecosystem surrounding EMF is skewed towards Java for historical reasons, but there have been some prototypes with other languages as well. You can serialize your AST in any way you like, although by default it relies on XMI files. You can implement your own textual concrete syntax, or rely on a database. The EMF ecosystem has tools for implementing textual or "graphical" concrete syntaxes. You can combine them (e.g. usually a specific subset of your AST gets edited in a certain way that's best for your targetted end users). The ecosystem also has tools for performing comparisons and plugging them into your editing means.
Of course all of this tooling requires a lot more work than an LSP server.
I think, the reason it was never implemented was that more translation = more complicated debugging. It also means that programmers have a more distorted and incomplete model of the program they are writing, i.e. more bugs.
NB. Lisp, as originally envisioned by McCarthy, had one more translation layer (the translated version had square brackets instead of the parenthesis), but it didn't take off for, basically, the same reason.
So... while I understand the benefits you see from doing what you suggest, I think that at the same time the downside makes this not worth pursuing.
.NET decompilers are common. I have built a few toy languages and compilers on .NET. For one of them, I could decompile CIL into my language. So, I could view .NET libraries from other sources in my language.
I think this is essentially the same idea you are proposing.
It only works if the languages are similar though. Going between F# and C# does not always work as well for example.
You are describing an entire industry of IDE Smell with an IDE monoculture.
Edit: I do agree and find your AST suggestion profound though!
I think we don't have any sort of flexible AST sort of thing because they're mostly not necessary. The hard problems of programming don't usually have much to do with syntax.
And if you're going to downvote, kindly explain why, thanks. I just want to know exactly how this thing is supposed to work...
I may recall incorrectly but AppleScript may be an example: some file formats are serialized ASTs. The editor displays it as textual code. A downside of this is that you can’t save a syntactically invalid file.
There are visual programming languages that chain together blocks, instead of raw text.
> The hard problems of programming don't usually have much to do with syntax.
I guess it depends on how you define hard. You are clearly talking about "a singular issue that needs to be solved", which really only effects a single developer / team and, to a lesser extent, those that use that solution. But if you consider something like syntax, you're now talking about something that much a much smaller impact _per developer_, but has that impact on _every_ developer. The syntax issue may have a much larger impact overall.
Sure it's an intuitive way of representing your data. Is it the most appropriate though? See an example [0] about using Projectional Editing in order to use mathematical notations for formulas.
[0]: http://voelter.de/data/pub/gemoc2014-voelterLisson-MPSNotati...
Here is an example of GPT's output for Python with braces that was generated after just spending 10 seconds for the prompt:
def preprocess_braces(code: str) -> str:
lines = code.split('\n')
processed_lines = []
indent_level = 0
indent_str = ' ' # 4 spaces for indentation
for line in lines:
stripped_line = line.strip()
# Check for opening brace
if stripped_line.endswith('{'):
processed_lines.append(indent_str * indent_level + stripped_line[:-1].strip() + ':')
indent_level += 1
# Check for closing brace
elif stripped_line == '}':
indent_level -= 1
else:
processed_lines.append(indent_str * indent_level + stripped_line)
return '\n'.join(processed_lines)
# Example usage:
code_with_braces = """
def example_function() {
if True {
print("Hello, world!")
}
for i in range(5) {
print(i)
}
}
"""
processed_code = preprocess_braces(code_with_braces)
exec(processed_code) # This will execute the transformed Python code
print("Processed Code:\n", processed_code)Isn't this the fundamental problem with code generated by a statistical text generation algorithm?
In other words, "code" is short for encoding a solution. And to have a solution to encode is to understand the problem to solve.
Without understanding, statistical code generation is little more than a popularity contest.
You can do a lot with ChatGPT or Claude if wrong code is easy to spot (which will obviously depend on what you're working on). If you can easily spot mistakes these things can often come up with a fix once you point it out. I've had some real success converting small-scale production C++ code into Python using Claude. Stuff that isn't really deep or complicated, but it's still faster and less annoying using an LLM to assist. I am sure there are large domains where it's crap, but for relatively simple stuff (CRUD logic, simple file parsing) it does remarkably well.
This is exactly my point.
Whether a programmer uses past experience exclusively to author source code or a statistical code generator (LLM) and then their learned ability is orthogonal to my original premise.
Without understanding, statistical code generation is
little more than a popularity contest.Code is not short for encoding a solution. Code is any kind of program whether correct or not.
You can have a solution without understanding it. You can for instance know the rough shape a solution should have and try to guess at the details. And get it correct some of the time.
And current models do have some form of understanding, although it is sometimes incomplete. They are clearly able to solve many problems after all.
In my case earlier today, it helped that it was a relatively simple function and I gave it a rather detailed natural language spec of what I wanted it to do. I totally could have written it all myself, but writing a natural language spec and getting GPT-4o to translate it to code is (depending on my mood) less mental effort than just writing the code directly.
% echo "nums 1 10 | filter even | to_words | map uppercase" | refab imagine
TWO
FOUR
SIX
EIGHT
TEN
% echo "with file '/tmp/top-ten-most-populous-cities.txt' do; cities = read; cities.each { |city| (city.name, city.utc_offset) }" | refab imagine
Tokyo, 9
Delhi, 5.5
Shanghai, 8
São Paulo, -3
Mumbai, 5.5
Mexico City, -6
Beijing, 8
Osaka, 9
Cairo, 2
New York, -5
For what it's worth, the tool isn't specialized for this, 'imagine' is just one of many prompts it can execute.Of course the execution is non deterministic and at the moment only works for simple things, but you can imagine as LLMs get more capable and more integrated with tools this will matter less and less.
I recently started writing a game in Godot. I don't know GodotScript, and I've found I don't like it very much in trying to learn. I turned to aider.chat to see if I could describe the functions, data structures, and systems I wanted and have it write them. I also tried writing in a more familiar language (...one with braces...) and having it translate those files.
It does pretty well, but it doesn't feel like software engineering. It's too hands-off and doesn't activate the same neurons. All the problem-solving and puzzle-solving is gone, and the successes are quite boring, and the failure modes are more irritating even if they're necessarily quicker to solve.
It's a weird experience. I'm moving so, so much faster than I would have on my own, but I don't enjoy it. It feels like cheating - I'm not actually ashamed of what I'm doing but I also won't take credit for writing the code.
However, what I'm getting at is this: If I could write the code in a syntax or even language that I prefer and have copilot or whatever translate it in near-real-time (without active prompting), that would be the best of both worlds. I'd still be a little sad at myself if I didn't learn the new language, but I also think this method would facilitate learning better than what I'm doing with aider (because I could see what my code turns into as I'm writing it, and learn that "translation").