Go 1.3+ Compiler Overhaul
docs.google.com
docs.google.com
For me the really fun bit is getting to having the whole runtime in Go, which would be a situation not dissimilar from classic Smalltalk environments - especially writing the GC in the target language.
Like the Jikes RVM? IBM Research did some really cool stuff. Anyways, anything is possible when you go meta.
Lets see if it really happens.
Porting the runtime to Go is also on the drawing board, but later. Right now the compiler complexity makes it harder to work on, hence porting it first is going to pay off quicker.
Sadly the GitHub Web UI seems to cap viewable history at 100 pages' worth of commits.
(The "Rust" that the OCaml compiler handled is very very different to modern Rust, fwiw.)
The OCaml compiler was just written to get an implementation of the language that was good enough in which to write a compiler. The OCaml source was deleted more than 2.5 years ago and Rust has been properly bootstrapping/dogfooding for that long.
So it's Go which still doesn't eat its own dogfood. I always wondered why Go was never functional enough to be used in a lot of the scenarios where C is used. With all the details, it's more clear now.
It's much better to be in the position now, when the language is barely used, than to have to support it later.
If it was today they would have done otherwise, based on current experience.
Well, at least Rosylin is going to be available in an upcoming .NET version.
"It is easier to write correct Go code than to write correct C code."
and
"It is easier to debug incorrect Go code than to debug incorrect C code."
and
"Go is much more fun to use than C."
Really are very subjective and don't add much beyond setting a POV for reading the article.
Subjective or not, since the whole project is about eliminating C from the codebase, they seem very relevant as well.
I evaluate the task and the language options. If I needed to prototype out a web service where the server was up to my control, I'd likely shy away from C/C++. In cases where I was writing a module, it would depend on the architecture/framework. Anything low-level or system related, C is my goto well above C++.
Stakeholders, external components, and team factor in.
I do love how simple and straight forward C is. Calling it hard to debug or write code correctly for is wholly based on the person making such statements.
The statements were that Go is easier to write and debug than C, not that C is hard. Considering that Go is memory safe and a significantly simpler language, it's hard to argue otherwise. As for the person making those statements, you should take a look at Russ Cox's resume.
> Stakeholders, external components, and team factor in.
Indeed they do. In this case, the stakeholders use Go, the external components are written in Go, and the team is the Go team. Go seems like a good fit.
But even without the context, your comments are in response to a misreading of a very small part of the overall document. Hardly seems worth your time.
Correct large-scale programs are very difficult to build and debug, full-stop. It isn't a "feature" unique to C, or any other language.
Likewise for correct code, and incorrect code: they're both actually trivial to do in any language in the small scale, and see above for the large scale.
That said, I view the effort to make a Go compiler using Go is something akin to "eating your own dogfood."
Also I think you severely underestimate the benefits of memory safety.
They have earned the right to "hubris". And, if they believe they've made a better, easier to debug, more fun language than C (and they do believe that, as it's clear based on their own discussions of the language), I don't think I'd feel any right to argue with them.
(The plan to transition over by automatic translation, not the bullet points up top, is the interesting part of the document to me. I hope they're able to achieve it without too-serious compiler-performance regressions, since I really like the zippy compilation I get now.)
> It is easier to write correct Go code than to write correct C code.
But the correct C code is already written.
> It is easier to debug incorrect Go code than to debug incorrect C code.
Is the C code incorrect? How much debugging is left to be done?
> Work on a Go compiler necessarily requires a good understanding of Go. Implementing the compiler in C adds an unnecessary second requirement.
The compiler is already implemented in C.
> Go makes parallel execution trivial compared to C.
Is it not trivial to run multiple C compilers in different processes?
Your arguments would hold a little weight maybe if there were no plans to continue Go development, but I doubt that is the case.
I think the design doc is talking about using shared-memory concurrency to do parallel codegen on a per-function basis (like some LLVM patches are experimenting with doing). In this regard it's easier to set up the infrastructure needed to farm out concurrent tasks in Go, because C has neither built-in channels nor the generics to conveniently build them.
(Go programmers like to write Go code, who'd have thought? :-)
[1] I think it was the scala compiler video rant recently published, or maybe SPJ on GHC.
You have to give the Go team credit for telling the truth in their last bullet point:
Go is much more fun to use than C
I actually think the other bullet points are correct. I just don't believe they are the reason for making the decision.
I just had an insight while reading this. It's a powerful insight. (I am surely not the first person to have this insight, but it's WAY ahead of being mainstream).
A programming language that could be PROGRAMMATICALLY REFACTORED would be a HUGE home run.
I know Lispers will jump in here. But I know Lisp (somewhat), and it does not have what I'm talking about. (Hell, maybe Go does ... if so, I finally get what the fuss is about).
The idea is that you could develop a huge codebase in this language, and then you could write code (probably in some other language) that REFACTORS the original codebase. Of course, you CAN write translators, as the original article mentions. But I want a whole new programming language that is designed from the GROUND UP to be amenable to programmatic transformation.
NOT just Lisp with its AST's and macros. The "transformation language" should be able to understand the following aspects of the original code: its modularity, its test coverage, which parts are functionally pure/impure, which parts are parallel/not, and the full compiler-level semantics of every piece of text in the code (in other words, what JetBrains knows ... this is a local variable name, this is a function name, etc.)
In other words, the original language has to capture more of the programmers intent (probably by inferring most of it). The intent-information is mostly or completely unneeded at runtime, but it is VITAL at automated translation time.
Imagine having such a codebase, and being able to pull up a REPL and interactively start changing the modularity of the code by issuing commands. Or telling the transformation system to parallelize some portion of code that wasn't parallelized previously. Or saying something like, "Take all the code snippets that instantiate the xyz data structure, and change them to call this function instead."
Don't miss my point -- we have features like this here and there. Some IDE's more than others, some languages more than others. Nothing new there. But I'm talking about a new programming language designed from the ground up to be highly amenable to this kind of interactive, automated transformation. In other words, I think this might be the killer feature that allows one programming language to outcompete most others.
This might be the one programming language feature that we should focus on now, more than any other.
Would any downvoters care to take a minute to share what you think I'm missing? Seems like the polite thing to do. I know you're busy, so make it short and sweet if you have to, but if I'm wrong/naive/ignorant, please, help me out.
What exactly do you think you're on to? Insight has built all of the things that you're hoping to replace. So what are you bringing to the table?
Who/what is "insight"? I didn't follow this sentence.
> So what are you bringing to the table?
If this has been done or thought out before, possibly nothing. But I'd be OK with that, because I'd be excited to see the results. If you have a link or a search term, please pass it on.
But if our future programming environments are "smaller and sleeker" in all respects(as is anticipated by, for example, the VPRI work on STEPS) this would be the wrong optimization to make. The cost of maintenance will go down across the board because the new languages let us express the change with less effort, and then complicated refactoring becomes less necessary again.
(Counterpoint: We just build even more complex systems and then need better refactoring tools. An endless cycle...)
And I get the irony! The exact same thing is true of my post. But I think I'm not trying to shoot as high as you are.
If we can find a way to make the VPRI/STEPS vision a reality, I'm all for it. I think what I'm proposing is a less ambitious interim step.
I didn't vote at all, but you state:
In other words, the original language has to capture more of the programmers intent (probably by inferring most of it).
But I think that your comment fails to provide (1) an argument why homoiconic languages such as Lisp and Prolog do not fit the bill, since they are easy to rewrite (code is data); and (2) how this hypothetical language differs over languages that are easy to parse and have simple semantics, such as Java and C#, if most of the intent needs to be inferred.
I'd rather like to see better libraries for e.g. Java (Groovy provides the REPL :)) to do this, than yet another language.
But the distinction I was trying to draw is that I'm talking about going further with that idea. Homoiconic syntax only allows for local code re-writing, but there are many other kinds of things we'd like to programmatically transform on a large codebase.
Half-baked examples: We'd like to replace all use of new/delete with smart pointers, but only in source files that are in one particular section of the codebase. Or we'd like to move all our unit tests from being in a separate package/dll, to being included in the same dll but surrounded by compiler pragmas.
(2) How the language would differ -- I don't yet know. But just as homoiconicity says "let's allow our end goal of automatic code rewriting to dictate our syntax", what I'm trying to say is, "let's allow our end goal of automatic refactorability / transformability to dictate all the programming language design decisions we make."
I think probably many of the things Raphael_Amiard mentions in this thread would be a part of it: strong inferred typing, perhaps, and a strong package system. Perhaps the ability to write some kind that was intentionally agnostic about how it would be packaged, and write some code elsewhere that says "take that code over there, and package it thusly." It would need to know if certain snippets were functionally pure (sort of how D does with "pure nothrow"), because then it would know it have more freedom with how it could transform that code.
Hopefully that clarifies the idea. It's hand wavy and high level, but it's a new (to me) idea that's only at the hand wavy stage so far.
For example, take your idea about wanting to be able to replace all instantiations with a function call instead. There's no need to invent a new language for this; the AST of a program in most any language would have enough information to write this refactoring capability into a tool. The hard part would be expressing the refactoring itself; how do you specify exactly the pattern you are trying to match, and exactly what you want to rewrite it to, in a way that is easy to use?
Maybe some other ideas you have would require a new language, but it's hard to say unless you are more specific about what you want. You might find that a lot of your ideas are implementable on top of existing systems. Designing "from the ground up" has a certain appeal to it, but to justify the extreme amount of effort a new language would take, I think you'd want to be more specific about what the clean break would buy you.
I agree that the AST of a program provides the raw material that could empower these kinds of transformations. But certain language features make it hard (like the C preprocessor, to cherry pick an example).
And I guess I'm exploring the question: What would a language look like whose goal was to make it easy, and to increase the set of possible automatic refactorings?
Maybe it's just a thought experiment that won't lead anywhere concrete. But to me, it's a very interesting thought experiment that I've never heard posed before.
I don't think it's a great idea. It's been tried in various ways (you assert that C# wasn't designed from the ground up to be automatically refactored, but actually easy manipulation of the code by IDEs was a key design goal - not the only goal to be sure, but a major one), and it doesn't seem to have worked. But don't let me stop you; try and build something useful out of this. I suspect you'll quickly see the problems with the idea, but if it actually works then we'll all be very grateful.
I think C# has more refactorability than most languages I've been exposed to. But I'm intrigued by the question of how much further could we go if we made that a primary goal.
Granted, it's just the merest kernel of an idea, and it's an extremely ambitious one at that. Not something I'm likely to tackle on my own -- more like a thought experiment at this point. Even if it failed, the attempt would be illuminating perhaps.
Anyway, thanks for leaving a comment.
But C# and Java weren't designed from the ground up around the primary goal of being amenable to such programmatic refactoring. So there are limits on what you can programmatically refactor.
But our industry has reached a point where "rearranging" large bases of working code might be a bigger, harder, and more important job than originally developing that code. So perhaps the most important characteristic of a programming language is its "refactorability."
(Update - I think this sums it up: Would you consider working in a clunkier programming language if you knew that it offered 10X more "refactorability" than the other languages you were considering?)
Anyway, thanks for the comment.
I'd have a hard time believing that a more clunky language would be a more refactorable language.
So, would you accept certain trade-offs in your programming language if it greatly magnified the programmatic refactorability of your code?
Also, have you tried Julia? It's homoiconic but looks more like Python. Haven't seen any refactoring tools for it though.
The go people are poking at this. They ship a tool, 'go fix', that can safely refactor your code to deal with backwards-incompatible API changes. Go also has a tool, 'go fmt', that most people use to enforce tabs vs spaces, etc., but has a 'rewrite' flag that you can use to do simple, arbitrary transformations.
They also ship some AST tools for go, though I haven't seen much use made of them.
Try loading this snippet, checking the "Import" box, and clicking "Format": http://play.golang.org/p/jS4s_Xz26v
To use it locally:
$ go get code.google.com/p/go.tools/cmd/goimportsWhen did that get added? That's brilliant!
And Slamhound is definitely barking up the same tree. Very interesting!
*The small problem with this example is that in fact the latter has been achieved.
Unfortunately, I freely admit that the joke goes right over my head. I used to know something about Riemann surfaces, so I could invest some time and eventually get the joke perhaps, but I just don't have time now. :)
Set Theory Natural Numbers Integers Rationals Real Numbers Complex Numbers Calculus Riemann Surfaces
(as the Apple ads say, some steps skipped)
The only point I was making was that the level you were thinking at was much higher than the level SEXPs are at. Trying to envisage how we get there remains hard. TL;DR I was agreeing with you. :)
That provides some incredible power and may be the thing about Go that impresses me the most.
What if we also made programming language decisions about syntax, modularity constructs, etc., all with that goal in mind? (And perhaps Go has to a large degree).
You should be able to specify a type name, identifier, or module and its' replacement (optionally file & line to be specific) via CLI. It should do the replacement automatically, and fail if it at all changes the semantics (e.g. clash or variable shadowing) and give an overridable warning on a non-semantics-changing possible problem (like an unused shadowed variable).
Sure, this is a normal feature for an IDE, but I feel like it should be simpler (and I don't like IDE bloat anyways). Maybe the tools you mentioned can already do this.
this is a pretty powerful package.
>(I am surely not the first person to have this insight, but it's WAY ahead of being mainstream).
Great. Maybe you can actually bring it into the mainstream. With code.
Imagination is fun, but you're not doing more than imagining.
He asked for real criticism, guys. His post is really high on imagination and low on delivery, and if you don't see that you should probably think about how many people have created our programming languages and contemplate doing better.
This person isn't having conversations with them.
Am I rude for requiring him to step up to his talk?
Edit:
>>In other words, I think this might be the killer feature that allows one programming language to outcompete most others.
This sentence alone is flagrant ignorance.
I can't step it up just yet -- all I have right now is the question, which I pose as a thought experiment.
This conversation is more akin to some physicists speculating on the limits of efficiency of solar cells, than some engineers talking about a concrete plan of action.
>>In other words, I think this might be the killer feature that allows one programming language to outcompete most others.
> This sentence alone is flagrant ignorance.
Not really ... it's just awkward wording on my part. I don't really believe this will lead to the One True Language to Rule Them All. I was just trying to say that it might be well worth sacrificing on certain language features in order to maximize refactorability.
So ... I don't agree with you, but I do appreciate you leaving a comment.
The idea behind all these type systems is really to allow automated transforms which preserve the integrity of the program.
Then there are languages whose evaluation is based on term rewriting, like the language Pure.
EDIT: You do have HLint, and the bot that rewrites expressions to point-free style.
I acknowledge that is not precisely what you were referring to, but it is close.
In my opinion, the problem is more of a tools problem than a language problem : One language that has one of the most interresting support for this kind of things is C, with coccinelle [1] [2].
Google has been working on absolutely amazing tools for C++ that applies programmatic refactorings in a distributed manner on absolutely enormous codebases [3].
This IMHO shows that this is a tools issue rather than a language issue. Coccinelle had to develop a swat of parsers for its project, and the Google project is using Clang as a basis for structural and semantic capabilities. C and C++, with their fragile type systems, preprocessor, horribly hard to parse syntax, might be the worse languages amongst typed languages to develop such projects on. And in spite of that, they are the ones for which such project exist, because they are the one with enough need for such tools.
Of course some languages are more amenable to such tools than others. Very strong static typing and a solid package system helps a lot. No metaprogramming helps too (because you don't have to handle the transformation).
I'm working on such tools for Ada, that is pretty much the perfect language for this as far as imperative languages go. The essential need for such tools is to have a compiler that exposes some services as an API, most notably the ability to explore the AST and query cross references for language entities. This was the fantastic insight of Clang/LLVM in my opinion.
[1] Coccinelle semantic patch language : http://lwn.net/Articles/315686/
[2] Presentation on refactoring with coccinelle http://video.rmll.info/videos/coccinelle-automated-refactori...
[3] Clang MapReduce -- Automatic C++ Refactoring at Google Scale http://www.youtube.com/watch?v=mVbDzTM21BQ
I'll be taking a close look at Coccinelle. Thanks for the link.
And as you say, C and C++ "might be the worse languages amongst typed languages to develop such projects on," due to things like the preprocessor, hard to parse syntax, etc. So I'm saying, perhaps it is time to invent the programming language that is most amenable to this.
It's amazing what we can do if we put our best engineers onto the hard problem of achieving this with C / C++ code. Imagine what might be possible if we designed a language from scratch with the primary goal of making this easy. As far as I know, no programming language has ever been designed with this as its primary goal (I'd love to be corrected if I'm wrong).
Transforming from one language to another is also old idea, some Scheme systems transforms (compiles) into C first.
In both cases the transformations themselves must be defined precisely by a programmer in advance. The idea that a program could transform semantics of another program is still a fantasy.) Even for Haskell.
I don't want to eliminate the role of human developers from the process. Instead, I want to provide tools that the human developer who understands these semantics can use to express the desired transformations, and tools that will help carry out those transformations across a huge codebase. Tools that "magnify" the efforts of the human who has the deep understanding of the before and after semantics.
It's almost as if they choose to tackle the easy problems instead of the hard ones.
However, I fully support this decision.
Bootstrapping is the best way to develop a language, as it allows the language designers to experience the language in first hand, while making it independent of other tooling.
Plus it is one argument less for C zealots against Go, in the sense of "my compiler compiles yours".
Generics might happen in Go 1 at some point. Exceptions will never happen. Explicit error handling was a design choice.
Pretty much any modern statically typed language can easily be refactored.
The problem is not the refactoring, it's deciding when to refactor and to what (e.g. renaming). This is much more of a human than a computer problem, and as such, it will probably be intractable for a very long time.
The averse reaction to goto is conditioned response, not learned experience.
It was already Pascal in academia and "regular" insdustry, C in low-level stuff and FORTRAN 77 in science stuff. Of course, there was still legacy Basic and legacy COBOL, and you were much more likely to run into a historical mess of GOTOs. But still, it was already past the Acceptable Spaghetti era.
In the annotator's introduction it stated something like "I haven't seen this many goto statements since I was writing BASIC as a teenager!" but then went on to explain why sometimes they're the right tool for the job.
I'll have to order another copy. It was an amazingly varied source of knowledge!
[1] http://www.amazon.com/Linux-Core-Kernel-Commentary-Edition/d...
A lot of people cargo-cult Dijkstra's hate on goto. It does have its uses as long as you're very selective about when to apply it.
We had an Actionscript 2 application that was under constant development and use on the web, and we needed/wanted to port it to Actionscript 3 to make use of new language features. I would write a tool that would attempt to translate between the two automatically, and then any more complex transformations could be hand coded and/or added to the translation engine.
Of course, the project did not succeed during my time at the company, mostly because the scope of the translator was so limited; there were too many concepts in AS2 (the borrowed javascript setInterval, for example) that had no equivalent line-for-line in AS3 (where you had to use the new Timer class to delay/loop execution) and therefore couldn't be simply substituted. Perhaps if the parser had tried to represent higher order concepts than files and lines of code (e.g. classes, instance data, methods) the project would have got further - but we simply didn't have the resource to spend on it, nor the ability to freeze features on the AS2 codebase.
As it was, the AS3 version of the app never quite reached feature parity with the AS2 version, which raced ahead too quickly for the new features to be added by hand in AS3 (and too many changes had already been made by hand to make re-running the parser viable.)
If Go is indeed simple enough to translate into from (a specific subset of) C, and if Google can afford more than one engineer for a week to write the tool, then this sounds like it's off to a better start than my project was.
Sounds like they had this idea in the back of their minds from the beginning, and wrote the C code accordingly.
The more likely explanation is that the same people wrote the compiler in C and designed Go. The subset of C they found most useful is the one they used in the compiler and the same one whose semantics ended up in Go.
In the second edition Plan 9 release, the last one with Alef, these programs were written in Alef: acme, aux/consolefs, aux/depend, httpd, md5sum, page (document viewer), postscript/tcpostio (driver for lp), and ppp.
9P was not so lucky (it's more of an OS thing than a language thing), but I wish it could take off somewhere.
Edit: to be more precise, they’d need to get the C syntax tree in some representation first. Then pass it over to a transformer. The latter could be in Go, but it seems foolhardy to reinvent the former.
Notation: T[A,B] : Translates from A to B
The OP wants to develop T[C, Go]. The question is which language should T[C,Go] be written in?
Option #1: If T[C,Go] is written in C, then it can use an existing C parsing library to parse C and then emit Go. Here parsing will be easier, but emitting Go will be laborious to write in C (since C is very low level).
Option #2: If T[C,Go] is written in Go, then the parent claims that C parsing will have to be freshly developed in Go. This will be a pain but the Go emitting code will be easier to write in Go.
However, I think the problem with option #2 (parsing) can be avoided if one uses the parsing library that is assumed to be already available in C. Since Go can interface with C code[1], the parser need not be reinvented.
Worse case scenario the resulting code will be a GOTO littered state-machine with 'import "unsafe"' on every source file.
Though I like to think the Go team would fare better then that.
With a translator work doesn't have to stop on the current compiler code base while the new version is developed, all that effort is saved and bugs are automatically fixed in both places so you don't have to maintain and synchronise two bug trackers.
Couldn't the compiler translation just be crowd sourced by a small group of people? It wouldn't be the best result but if people just stuck to the original C design, it would work as a first implementation.
Meanwhile, development on the C++ implementation continued. This is probably the biggest advantage, as the porting does not slow down language evolution.
[0] http://forum.dlang.org/post/l8uccc$1qkr$1@digitalmars.com
Looking forward to the day Go's runtime is free from C code.
And I'm not sure Go qualifies as a "general purpose language".
They plan to keep around the C based compiler for the time being until the whole process is finished.
Afterwards there is the option to have a backed that generates C code as target, although other approaches can be taken as well.
They just need to replace that backend by other one that cross compiles.
C is nothing special, I already wrote a few compilers without a single line of C.
I feel somewhat foolish for asking, but how do you define a "general purpose language"? I've seen Go used for building internet servers, command-line tools, games, scientific computing, mass data processing, and more. Sure, there are some things Go is not well-suited to, but I think it does enough (and well) to qualify as "general purpose".
The D compiler is waaaaay better, written by a compiler expert.