Porting the Go compiler from C to Go
gophercon.sourcegraph.com
gophercon.sourcegraph.com
Dang it I cheated and looked up Rus's quine. For anyone not wanting to cheat a good hint would be "tail recursion"
A little dream of mine would be if in the future when Rust is stable Mozilla developed a transpiler from C++ to Rust. That would be brilliant.
By the way, all the other talks at GopherCon seem pretty interesting, I hope someone uploads videos of them soon.
C++ to Rust would be even crazier.
(Transpile is a silly word. It's "compile". Compiling already trans-es.)
But maybe a set of guidelines about translatable C code plus a translator which is good enough for a variety of cases would make refactoring the resulting code such a manageable task that many projects could consider switching to Go directly.
You are technically correct in that compile already implicates a translation, but we usually use that term refering to a translation into a lower level language and not another language at the same level. You can say transpiler or source-to-source compiler for those cases and I think it's clearer and more accurate for the reader.
Rust is not lower-level than C++ and neither is Go lower-level than C.
No, he said a compiler translates from a higher-level language to a lower-level language:
"You are technically correct in that compile already implicates a translation, but we usually use that term [= compile] refering to a translation into a lower level language"
The way I understood what he wrote, he was making the same point as you. Presumably the people downvoting you had a similar interpretation of what he wrote.
However, arguably the term could make more sense there, if we ignore the earlier definition and instead assume "trans" is short for "transcendent", as in "climbing to a higher level".
If we then simplify that to "compile from one language to another of roughly equal or higher level", it becomes a useful word to indicate a specific subset of ways one can do compilation.
Of course, I'm completely pulling this out of my ass and I might piss off actual CompSci people who use very specific and strict definitions in their papers (kind of like how some get annoyed with the procedure/function mixup), so take this with a grain of salt.
What comes out the other end, as Russ phrases it, is a C program in Go syntax. The next phase of work is to turn it into “Go” as a human might write it.
The translated program is not meant for deployment, really. It’s to give the humans a starting point.
I think the best approach for something like this is a disassembler for LLVM IR. I am not extremely familiar with the LLVM IR, but I figure there are common patterns or debug symbols clang produces that could let you get a decent idea of what the higher level gave you. If not targeted to clang output too much, you can then transpile any LLVM-able language to Rust.
I think, but am not sure, that the problem with Rust being a target for these other languages (my pipe dream is to write a JVM in it w/ a JIT using the bootstrapped rustc lib) is that you have to use unsafe code all the time (unless there's 0 performance penalty for ARC) because the original code is not written with ownership/borrowing semantics.
They want to go quite a step up from doing some global regex replaces, but in the end, step 3 has "cleaning up and documenting the code, and adding unit tests as appropriate.". I don't think they aim to automate anything there.
https://docs.google.com/document/d/1P3BLR31VA8cvLJLfMibSuTdw...
I'm curious how they will handle things like pointer arithmetic and memory safety in C vs Go. If they mange to do so in a performant way I could see translating lots of numerical or computationally intensive code to Go so that it could be run in a shared cloud environment without worries about memory safety and without having to resort to vms for separation.
For a project I had (in the early 90s iirc), I had to extensively modify the fortran and make multiple versions before the generated C was readable. Still, much easier than a rewrite.
Wow, that's really striking. I know that goto statements have their uses but for something written in the last couple of years to have over a thousand of them is very surprising (it might not be surprising at all for those who write C code all the time). I guess they're mostly just for error handlers?
Some are for single error return, but others are just to make a nice structure for the code.
This smells of very old code (as the header testifies)
> Portions Copyright © 2009 The Go Authors. All rights reserved
That's pretty recent.
And here's BSD style(9): http://www.freebsd.org/cgi/man.cgi?query=style&sektion=9
The function type should be on a line by itself preceding the function. int
main
(int argc, char **argv)
{
//body..
}
although most people think i'm weird for it.. int main(int argc, char **argv) {
//body..
}
I thought it was nice but not obviously advantageous, but over the years I've come around to thinking it makes code easier to follow. Even many bash shell scripting guides now recommend a similar form: for x in y; do
# body
doneIt is also a pretty common style in the open source world.
> It is also a pretty common style in the open source world.
Depends, one proeminent C project doesn't use that: the linux kernel
And I can do it without ctags. That's the point.
>Depends, one proeminent C project doesn't use that: the linux kernel
That doesn't make his statement "depends" at all. A single project not using that style does not contradict it being "pretty common".
I do wish that in C1x (for x > 1), they could find a way to let us declare multiple formal arguments of the same type without repeating the type:
float
my_graphics_hack(float x, y, z, r, g, b, u, v) { ...
instead of float my_graphics_hack(float x, float y, float z, float r, float g, float b, float u, float v) { ...
which gets really tedious and obfuscates that fact that they are all the same type, especially with the typename is more complex than just "float". When the new syntax was added to ANSI C, it wasn't possible to do multiple variables, for reasons I've forgotten but which I think had to do with forward type references. It would be awfully nice to find a fix for this.Preferring that function names start at the first column to make searching easier is a perfectly good justification.
You, on the other hand, seemingly would impose your tools (which have their shortcomings) and workflow on people with no justification at all.
Just because a tool exists, doesn't mean it's good (or better than what people are accustomed to) or that everyone must use it.
Do you think everyone should use Vim on Linux too?
Adding syntactic sugar to a language makes the language bigger and harder to completely understand, it makes it easier to misuse a feature (and C already has a lot of trickery with types, like when you declare a parameter as an array but it behaves like a pointer). I'm a strong believer that implicit is better than explicit and that while there are many ways to do the same thing some are better than other in practice. Of course, "in practice" can change from one project to the other, what matters in consistency. I applaud the choice of Go to standardize on a single coding style for instance.
For the particular example of the parent, while I write a lot of C code for work and for fun I have yet to encounter a situation where I saw a function taking 10 floats as arguments and thinking "yup, that's completely the right way to do that". If you have an example of such a code I'd me more that willing to reconsider my position, otherwise we're just talking about the best way to tame a unicorn.
You can also search for something like "\w+\s(.?)\s*{"
It's not about finding the function name in general, it's about finding where the function is actually declared -- not in the call sites.
>You can also search for something like "\w+\s(.?)\s{"*
Yeah, or use ctags. That's why they said "simpler".
"^foo" beats that.
> "^foo" beats that
because it's so hard to use ctags...
grep -n ^functionname *.c
and get the line that starts the definitionIf you're writing a recursive descent parser for C while keeping minimal lookahead, I can think of a place or two where goto might really help. One is to deal with casts, compound literals and subexpressions. They all begin with a left-parenthesis, so you can't tell what you're dealing with until a bit further. Since these all appear in the parsing of simple expressions (identifiers, constants, unary operators), the code can go in one function and you can jump to the right place with goto as soon as you know what you are dealing with.
I did not look in the Go code but C programmers mostly use goto for error paths.
"goto considered harmful" is taken out of context and bemoaned mostly by people unfamiliar with idiomatic C, I think. (Edit: sorry, I mean just for C. For other languages, e.g. Go, there may be more appropriate patterns like the 'defer' keyword.)
I don't think the authors were very goto-averse.
Hardly surprising. Gotos will be found by the hundrends or thousands in C projects like compilers, kernels etc.
It's used for localized consolidated error handling, but also for stuff like parsing code (bison generates tons of gotos IIRC), state machines, etc.
If anything it's the old "Goto statement considered harmful" that's a little too naive.
No, people who didn't read it and just repeat the title without knowing what he was talking about may be too naive, but the original point was not.
I understand the desire to promote the Sourcegraph app by doing the blogging, and I think its effective. However, the blog is real annoying to browse, as every (prominent?) link points to Sourcegraph the app instead of the blog.
I'm REALLY thankful for sourcegraph's liveblogging. The above comment was a suggestion as I thought they might want to know that (at least for me), the navigation of the blog was confusing.
When transcribing a talk, there isn't any need to write "They're". Just use the same pronoun the presenter used, otherwise it stands out like a sore thumb.
In what sort of ways does self-hosting early influence a language design? Were they hoping to avoid something in particular by delaying self-hosting?
The part I quoted above almost sounds like wiping sweat from one's brow after having dodged a bullet: "phew, the language is safe from influence by those gull-durn compiler-writers..."
Anyone have any idea what the title of that book is?
>A Union is like a struct, but you’re only supposed to use one value
>(they all occupy the same space in memory). It’s up to the programmer to know which variable to use.
> There’s a joke in some of the original C code:
> #define struct union /* Great space saver */
> This inspired a solution:
> #define union struct /* keeps code correct, just wastes some space */
Somewhere in Scotland, a sum type sheds a single tear. > #define union struct /* keeps code correct, just wastes some space */
Not always, though...Probably about the only place you'll see them used like that is in the code produced by web2c in compiling TeX. Knuth used variant records a lot to get around Pascal's type safety, and they get translated to unions.
#define union struct
Here you go: #include <stdio.h>
union U { char a; char b; };
int main() {
U u;
u.a = 13;
u.b = 10;
printf("%d\n", int(u.a));
return 0;
} struct A {char a;};
struct B {char a;};
union U {struct A a; struct B b;};The only standard compliant way to, say, convert a float to an int is to use memmove:
uint32 i;
float32 f;
i = 0x80000000;
memmove(&f, &i, 4); typedef union
{
#ifdef TeX
glueratio gr;
twohalves hh;
#else
twohalves hhfield;
#endif
#ifdef XeTeX
voidpointer ptr;
#endif
#ifdef WORDS_BIGENDIAN
integer cint;
fourquarters qqqq;
#else /* not WORDS_BIGENDIAN */
struct
{
#if defined (TeX) && !defined (SMALLTeX) || defined (MF) && !defined (SMALLMF) || defined (MP) && !defined (SMALLMP)
halfword junk;
#endif /* big {TeX,MF,MP} */
integer CINT;
} u;
struct
{
#ifndef XeTeX
#if defined (TeX) && !defined (SMALLTeX) || defined (MF) && !defined (SMALLMF) || defined (MP) && !defined (SMALLMP)
halfword junk;
#endif /* big {TeX,MF,MP} */
#endif
fourquarters QQQQ;
} v;
#endif /* not WORDS_BIGENDIAN */
} memoryword;
Once you sort through all the ifdefs and typedefs, it boils down to a union of a float (glueratio), an int32 (integer), a pair of int16s (twohalves), and a struct of 4 bytes (fourquarters). When TeX creates its memory dump file, it does it all in terms of the bytes in the fourquarters struct. (And I think there are other times that it does similar things.)This may very well be undefined behavior, but it seems to work when compiled with GCC anyway. Thankfully gc isn't such crazy code as this.
``One special guarantee is made in order to simplify the use of unions: if a union contains several structures that share a common initial sequence (see below), and if the union object currently contains one of these structures, it is permitted to inspect the common initial part of any of them anywhere that a declaration of the complete type of the union is visible. Two structures share a common initial sequence if corresponding members have compatible types (and, for bit-fields, the same widths) for a sequence of one or more initial members.''
union hack { int x; float y};
int foo( int *i, float *f)
{
*i = 3;
*f = 42.0;
return *i;
}
Aliasing rules say foo need not read from that int pointer, may switch the order of the write to f and the write to i, and can assume that foo returns 3, so struct hack h;
foo( &h.i, &h.f);
might return 3 or something else. I think that last call introduces undefined behavior, but only becuase of the definition of foo that the writer of that call might not even have the source for.But of course, that is an "you shouldn't do that" edge case. One could also claim that the corrigendum doesn't apply because foo doesn't "use a member to access the contents of a union".
And I agree that the corrigendum doesn't apply in this case. Once you hand different-typed pointers to the same memory to people, whether via union or just casting pointers, the aliasing rules will up and bite you.
typedef union
{
struct sockaddr sa;
struct sockaddr_in sin;
struct sockaddr_in6 sin6;
} addr__u;
as a way to work around the BSD socket API without horrible casting all over the place. Change the "union" to a "struct" and the code no longer works. Is it subverting the type system? Yes. Is casting subverting the type system. Yes. Pick your poison.