Tor in a safer language: Network team update from Amsterdam
lists.torproject.org
lists.torproject.org
It feels like it was invented in a universe where Haskell, OCaml, Erlang, Smalltalk, Lisp and so many more languages and research in languages never happened.
You can pick it up in a weekend. A lower entry bar means more people will try it out.
> It feels like it was invented in a universe where Haskell, OCaml, Erlang, Smalltalk, Lisp and so many more languages and research in languages never happened.
It was developed in a large enterprise context not in an academic context. I think it shows.
Which makes Go's rapid popularity as a language for solving the same problem even more peculiar. Especially since Go's surrounding tracing, debugging, and online code swapping facilities are so much worse.
I'm sure some component of that is Ericsson's complete lack of interest in evangelism and Erlang Solution's apparent lack of capability to effectively evangelize, but it still seems like Go isn't so much filling a void as it is being better at attracting an ecosystem that delights in reinventing a particular set of wheels.
Java was very much the same way for a very long time.
I'm probably outing myself as the lord regent of all impostor plebs, but Erlang is not IMO an approachable language regardless of its origins. Put differently : it looks completely nuts. I'm sure there's a method to the madness, but comparing it to Golang feels misguided (at least, from the perspective of awful programmers like myself)
Elixir by comparison is significantly more complicated, but the fact that it looks tacitly like Ruby is apparently very attractive.
I suspect I could sit down with you for an hour and remove all the weirdest seeming bits which feel alien. Such that the rest of it would just feel like any other programming language, and you'd be able to immediately start dissecting and understanding other people's code.
Part of the problem with Erlang (again an evangelism problem) is that there's not much in the way of attractive and we'll structured on-boarding documentation. It's mostly really simple. It's just that the Erlang docs seem to go to great lengths to obfuscate that fact. In large part because they don't do anything to intentionally build "background schema" with the reader by drawing comparisons to similar things they're probably already familiar with in other ecosystems.
{ok, u_no_me, {down_with, init, [State, Mod, OtherMod, Pid, SomeImportantRefIforget]}, {permanent, brutal_kill, 10000, 9999, 100000, 5, worker}
Is not a genius language construct of a faraway Alien race. It's not elite. It's not even Swedish. It's just limited and unclear. (there are even more unclear examples in the annals of Erlang, but everyone is familiar with consulting and reconsulting the child specification documentation)There are just so many things that make dealing with OTP nicer in Elixir. This is to say nothing of the meta-patterns that Elixir has driven home: Agents, Tasks, the actual supervisor/worker interfaces themselves. You could go on here.
I haven't even gotten to macros -- Have you ever had to make your own behavior, say for an acceptor/worker pool of some variety (presumably non-tcp, otherwise just ranch that out). Making your own behavior (what armstrong himself called "really advanced") is the easy part here. You've got to create a system where you pass around the TRUE module that holds the correct `handle_call/cast/info` implementation back and forth through the true module and the OTP state parameter.
This whole pattern (which comes with very real runtime costs) just doesn't even exist in Elixir. You just create a macro. A macro that sets all of this up at build time. No complex supervisor hierarchies where you need to constantly inspect each and every message, no custom behaviors, etc. You just have a mostly sane way of sharing software logic to begin with.
You can work up really clever solutions with parse_transforms, even standard "-define", but today Elixir has a totally brilliant solution to this problem and many more classes of problems.
This isn't to say there an't blindspots (beyond "wrap proc_lib/inets/other hated library"). There's minor issues here and there. Process registry naming can get confusing, especially operating between Erlang+Elixir supervision trees. Low-level details aren't nearly as well-publicized in Elixir either. There are even big issues, like the Elixir community leaning more "I'm playing around with Elixir because Phoenix is Webscale" than the grizzled realtime adtech/gambling/finance gurus and so if you want help with OTP platform/VM details, you going to have a harder time.
I first tried Elixir in v0.08. & function capture syntax had just been standardized. I suspect that like me, you tried it then, when it wasn't a fully-featured alternative yet. Things have changed and Jose Valim has a vision for the future of Erlang that I think you'd be foolish not to pay attention to.
Having used both Erlang and Elixir "in anger" there are a whole bunch of things that bug me about Erlang. There are a whole bunch of things that bug me about Elixir too. I'm not sure what any of that has to do with what I said. Elixir is a more complicated and larger language than Erlang.
It has structs (which are weirdly implemented as maps instead of records). It has type-classes/traits by way of "protocols" (which are a whole different layer of metaprogramming that could be replaced with behaviours and data-structure definition convention or dialyzer type-spec'ing). It has the pipe operator (which strangely elides the first argument in a function such that map/2 becomes map/1 when written... why no placeholder sigil). It has Agents (which are something which is confusing to have in the standard library as its a specific implementation of a gen_server which acts as a KV-store which will be really, really prone to data races). It has a fair bit of metaprogramming to be aware of in terms of macros and hooks that trigger at code loading/import time (which forces one to need to be a lot more aware of the inner details of 3rd party libraries being used). It eventually devolves into requiring knowledge of Erlang when attempting to do anything regarding debugging, tracing, or distribution related.
Elixir is a great language, and it has a lot of neat features, and I love teaching it to people and being increasingly critical of Erlang based on advancements in usability and community engagement/feedback that I see happening around Elixir. But I don't really need to be proselytized to about it.
* Inasmuch as you can make a blanket statement like this about a language, Go's speed is very much comparable to Java's.
* Java has a huge number of garbage collectors available. You're just talking about the default, but there are extremely low-latency GCs.
* Java has a bunch of ahead-of-time compilers; thanks to Android this might be the most common deployment of Java.
Even green threads were tried (and sensibly abandoned) in Java before.
I think people may not be aware of all the options available in the Java ecosystem, but other than sized types, which are theoretically coming to Java 10, there isn't much performance-wise that Go does and Java doesn't.
1x time: C
1.5x time: C++ (with smart pointers)
~2x time: Objective-C (almost everything is a heap smart pointer)
3x time: Java, Golang (optimized GC languages)
Session establishment for PSTN can be relatively expensive (much in the way of TLS handshakes), so concurrency and shared-nothing memory model together allowed for real-time streaming to keep on working no matter what else happened. The three main features of a PSTN switch are, after all:
1. Reliable call switching
2. Reliable real-time throughput
3. Reliable billing and accounting data generation
We don't think much of throughput these days, when any home office switch has gigabit ports and 40Gb+ backplane. As far as I know, maximising throughput bandwidth was not a primary consideration with Erlang. Reliable real-time streaming is much more about guaranteed latency - and incidentally, optimising between latency and throughput tends to be all about tradeoffs.
> You can pick it up in a weekend. A lower entry bar means more people will try it out.
... and then you get the JavaScript situation where everyone thinks they're an actual programmer at $TEAM_LEAD level of sophistication when it comes to modeling, design, architecture, testing and implementation... but they're really not. It's not their fault per se, it's just that they do not have the necessary experience to realize the areas that they're lacking in. (This is a well-known cognitive bias: More than 50% of people think that they're an above-average driver. There are a precious few who realize that they're not -- they are the exception.)
Programming is hard[1] and if there was an instant-humbling device, I'd buy it in spades.
[1] Most people focus on the "oh, that's Undefined Behavior" bits of it, but it's not really about that. It's about recognizing the types of mistakes that you make and working to avoid them or creating a system that does, whether that's by testing or proof or whatever. Obviously it still has to be _practical_, but AGAIN... it's about tradeoffs. If you don't care about correctness, I'm pretty sure I can whip up a "solution" to any problem that's fast as hell... and not correct. Summa summarum: I think we need to be thinking a lot more in terms of tradeoffs and not so much in absolutes.
I've seen a lot of people in academia writing C++, mostly horrible code. I've seen electrical engineers writing awful assembly code. Because the incentive is mostly to get one task done today. None of these languages have a lower entry bar.
JavaScript is pretty much bound to a singular use case (web development) where there is a lot of incentive to get quick money, which attracts all kinds of people. I don't see the status quo you're referring would be much different if the browser scripting language was Haskell or brainfuck.
I agree, but there are a lot of jobs where it's actually effectively impossible to distinguish between good vs. bad practitioners of said jobs.
I just have this feeling that it should be possible in programming. (Because it's quantifiable... or at least quantizable into "works" or "doesn't work" along any number of axes.)
> I've seen a lot of people in academia writing C++, mostly horrible code. I've seen electrical engineers writing awful assembly code. Because the incentive is mostly to get one task done today. None of these languages have a lower entry bar.
Same here, only with FORTRAN and C.
(This hints that this may be a meta-problem.)
> JavaScript is pretty much bound to a singular use case (web development) where there is a lot of incentive to get quick money, which attracts all kinds of people. I don't see the status quo you're referring would be much different if the browser scripting language was Haskell or brainfuck.
Well, there is node.js, but really, the problem with JS is... JS. It's just a horrific language semantics-wise (syntax: meh). It was invented and implemented in ~7 days(?) and it really shows. (No blame towards Brendan Eich, he did the best that he could within the deadline and even got a little bit of higher-order programming in there.) It's been improving, but just imagine the burden of improving a language that's already been deployed to 1B+ computers. Not an enviable task.
Hopefully WebAssembly will (in time) address these problems and give us a real way to program the front end in $WHATEVER_LANGUAGE_YOU_WANT
I think Pike and the other designers skew more towards corporate research (Bell Labs). And surely the development of GO as well as other Google research projects are intended to win in the market place.
From https://talks.golang.org/2012/splash.article > The Go programming language was conceived in late 2007 as an answer to some of the problems we were seeing developing software infrastructure at Google
I know it sounds like a lame reason to dislike a language, but I've always found staring at Lisp to be so much more difficult and distracting than C-style syntax.
printf("Hello world");
Lisp mode, move left parenthesis, remove semicolon (printf "Hello world")
Second round if (var == 2) {
do_something1 ();
} else {
do_something2 ();
}
Lisp mode, move left parenthesis, remove semicolon and curly brackets (if (= var 2)
(do_something1)
(do_something2)
)
On a big source file I am not sure if the amount of parenthesis is bigger than parenthesis , bracket, curly brackets and semicolons counted together.And then you discover paredit and you start wishing every other language would let you treat your source code like that.
It's just never been a go to language for me.
The latter group tend to find lisp not that painful because lisp parens/s-exps are as explicit as you can get, and indentation makes parens almost invisible
https://en.wikipedia.org/wiki/Dylan_(programming_language)
Also, Julia was syntactic and/or semantic sugar on top of femtolisp. Got converted to it in first pass with LISP's power doing the rest. Here's Stefan Karpinski on that:
"So ultimately the reasons for Femtolisp are:
1. Scheme is excellent for writing parsers since trees (aka S-expressions) are its forte.
2. Femtolisp is a small, simple, highly embeddable and remarkably fast Scheme.
3. We control it (and by "we" I mean Jeff) and can fix any bugs we encounter."
Yep, No 1 are those godawful parentheses and s-expressions making the job easier. ;)
One of my professors used to tell us, that C was built by people who wanted to use it and didn't care about academic style.
In many ways Go is just the next step of C. C did not have object orientation and even passing functions around was kinda hard. While C++ tried to bring object orientation to C (total failure) Go decided to keep the core values of C and instead improved the rough features (e.g. easier binding of functions to structures, faster build times).
By making it easier to pass around functions Go enables functional programming styles, but at its core, it is still just an improved C. The only revolution within Go (as a language) is the concurrency and channel concept and I think that was taken from some functional language (not sure).
I really like Go, because it just feels right. It might not be as clean as Smalltalk or Lisp, but it has data structures and functions, teaches you how important interfaces are and lets you build highly concurrent applications with ease. In addition, it brings a nice set of tools which integrate well into a shell driven workflow.
After all, the whole thing should not surprise anybody as Ken Thompson[1] was part of the Team which invented Go.
I'm not a Go programmer, but I have a lot of respect for it.
If you're wondering "why Go"? Think of it as a modern version of C, at a slightly higher level, developed by the same people for slightly higher-level tasks. They made C for low-level stuff, and then picked up with Go for higher-level stuff. It's like C+.
C has been very successful in part because it's so simple in certain respects (although not in others) and I think Go will be successful for many of the same reasons. Go does what it does very well.
I think Rust is actually a great choice for something like Tor, but I wouldn't use Rust for some of the things I'd use Go for.
That's perfect.
The one improvement over C related to build time in Go is that you don't need include files. Granted, that's a big one. But modern C compilers make header parsing extremely fast.
[0] https://en.wikipedia.org/wiki/Communicating_sequential_proce...
I think we've seen a few posts like that on HN and now I wonder, what that means exactly. What else falls under the "feels right" umbrella for you?
- In Scheme/Lisp, you write the function name before the opening parenthesis. In languages related to C, you write the function name before. For me, the second one feels better. In general, I like the C syntax pretty much.
- In C++ you easily provoke very long compiler error messages. Go compiler messages are much shorter and much more to the point. I like the ones where I do not have to search the real error in the error messages themselves. (I heard Rust compiler error messages are even better).
- To run a go program on a computer you just need the binary, done. To compile a program you need the compiler which comes with a few cli tools, done. For java, you have to decide if you need the JRE or JDK, agree to some license before being allowed to download it from their website. In addition, you have to place the jre on every computer which should run your program. I like the simple way.
- When I design a program there are a few parts. One is the Entity Relationship Model. Sometimes I think about it as something that has to be saved to an SQL database, sometimes as an object oriented inheritance hierarchy. For both there are reasons, but even simpler is it to just think about it as a struct or JSON. This fits pretty good for designing JS and go apps.
- When I write bash script I know what a slow language feels like. I know there are faster languages than go (e.g. fortran), but I am also aware that much of the performance is in the hand of the developer (e.g. memory management). When I write go programs I feel like my efforts to make the program efficient are worth my time and the tools support me while doing so.
- Last but not least go supports then functional programming style when needed. I like that. Sometimes I miss object orientation a little, but that mostly happens when I forget what interfaces are for. And while I truly respect Alan Kay I think object orientation has been overused/misused enough, so that its ok for me, that go didn't build it into the language.
So maybe I just like 'simple'. I am sure, that for every example I listed you can find a language which does even better than go, but I think it becomes harder when you consider all examples together. Nonetheless, the list is far from complete, I just tried to write down out of my head, what I like about go.
after
> the second one feels better
Not really. Parentheses makes code manipulation easy. Thus it actually it FEELS better, though it may not LOOK better.
Nevertheless, please elaborate. What do YOU think is the difference here between feeling and looking? Do you mean that your favorite editor (or the majority of editors) better supports outside parenthesis?
If the cursor is here
(foo bar)
^
I can cut the whole thing using just d%: combine d)eletion with the % cursor movement (jump to opposite parenthesis). Then you can paste it elsewhere with p.It's a POSIX-standard feature: see here: http://pubs.opengroup.org/onlinepubs/9699919799/utilities/vi...
That's just a crap editor for sysadmin tasks I wouldn't use for development.
If the syntax is:
foo(bar)
^
then we jump only over the argument list, not the whole expression.Another thing is the damned comma disesase in f(x, y) languages. Say we are in Vi:
(foo abc def ghi)
^
We want to swap the last two parameters. Easy: type deep. Done! d)elete to e)nd of word, go to e)nd of word, p)aste.Now try it with
foo(abc, def, ghi)
^
Annoying!Imagine if your operating system shell forced you to use commas between command line arguments. Nobody would use such an idiotic thing. Why do we put up with languages that do that?
$ ls, -l, *.foo # just kill me now
Move second argument to third position: "parens outside" together with "no commas between arguments" makes it a breeze: (foo (a b c) (d e f) (g h i) (j k))
^
Instead of deep we just do d%%p. Done. (foo (a b c) (g h i) (j k))
^ d%
(foo (a b c) (g h i) (j k))
^ d%%
(foo (a b c) (g h i) (d e f) (j k))
^ d%%pI think I will still prefer the C style for parenthesis (as I would care more for the LOOK than for the FEEL ;-), but regarding the commas, I agree that skipping them would make the life a easier. I mean besides lisp and bash/shell other languages have done so too (e.g. Smalltalk) and spaces aren't allowed anyway inside parameter names.
Yes.
> Nevertheless, please elaborate. What do YOU think is the difference here between feeling and looking? Do you mean that your favorite editor (or the majority of editors) better supports outside parenthesis?
Lisp has a two-level syntax. The first level is the syntax of s-expressions. On top of s-expressions we have the actual Lisp syntax.
S-expressions have a few features:
* it's a data syntax for lists/trees, numbers, characters, symbols, strings, ...
* delimiter surround the data, it is always clear where the expression begins and where it ends
* whitespace is used to delimit the elements
* s-expressions are not sensitive to lines and whitespace
* s-expressions can be automatically formatted by simple rules, according to different widths
* the tree structure is explicit, not implicit. It is visible, based on the s-expression nesting.
For an s-expression editor it makes not much difference to edit a data list like ((berlin germany) (rome italy) (paris france)) or code like (defun collide (object wall) ...) .The first level of editor support you get for editing s-expressions.
Thus editing on this level FEELS like you manipulate data: create, transpose, delete, list, de-list, flatten, copy, indent, format, ...
That every list has explicit delimiters makes clear where the expression begins, where it ends and what its contents are. The parentheses also serve as 'handles' for the thing. If you use some more advanced Lisp system, the s-expression creates a region and moving the cursor into this region enables context sensitive commands. This is possible in other systems, too. But here the relationship between the s-expression and the region is visually clear: each expression has explicit delimiters, front and end.
So, the first level of Lisp editing is data manipulation. That's a big difference to editing many other languages, where your program is not also a simple data-structure. There you are always on a language level, maybe on a primitive token-scanner level. You can reconstruct the tree structure, but it is not visible, explicit and delimiting like in Lisp. If you refactor a program, you work on the programming level - in Lisp you can work on a plain data level, too. This makes code and data interchangeable and when you work with a Lisp listener (running a read eval print loop), you will work with code as data and the listener helps you: you get support on the language level & the s-expression level on the editor side. But at the same time you can cross the the border into the programming language: you can let Lisp manipulate your program. Thus programming becomes a mix of manipulating text and data. The s-expression syntax helps to make that simple - because of the features above.
A typical example would be writing a macro (Lisp code which transforms code) based on some existing expressions. You would take the expressions, convert them into data, create the transformation code, define the macro. Then you would test the code generator. Thus suddenly from writing code, you switch to writing code-writing-code and the input and output is no longer data, but code as data. Thus while programming you will interact with the code generator. This can be done in many languages, but in Lisp it FEELS different, because you work on s-expressions - easily delimited hierarchical pieces of code as data, which can be transformed by your editor and your underlying Lisp system.
After a while, editing conventional code FEELS less direct. It feels like you manipulate the code with instruments, while a good Lisp system feels direct. That's it direct manipulation of code and data.
A Lisp programming will learn this code manipulation side and then is willing to give up some looks for that. Originally Lisp had a more traditional surface syntax, but it turned out to be more practical to use the s-expression based code representation not only internally, but also using it externally on the display or textual side.
Some time ago I used DrScheme (now Racket) to write Scheme but I never found it to fit easily into my workflow/use-cases.
So I would I like to use the language with the following workflow (100% terminal):
- write code with vim: vim main.lisp - simply compile the code to binary: lisp build main.lisp - execute the compiled program: ./main
Any ideas?
Essentially for anything slightly complex one would use an interactive programming style and 'only' deliver the application in a batch style.
Racket is more oriented towards batch programming, compared to popular Common Lisp development environments.
You can edit a Lisp program strictly in files, and run a build step to produce a clean image which is tested, using a REPL just as a debugger, to "go in" and find out what is wrong.
Still, Lisp is good for that. I've debugged Lisp programs with print statements and it was at least as good an experience as debugging programs in other languages using print statements.
For instance, if we compare to C, C has no trace, and generally no easy way to wrap any function with a wrapper that takes the arguments. Ecosystems built around C, like the Linux kernel, have developed things like that: Linux has a function tracing thing in it (more than one, I think).
So you use a language that makes specifying ownership (and immutability hard)? I always feel Go adds a lot of mental overhead.
E.g., if you have a method that returns, say, a float32 slice. 1. If I return a slice of an array/slice that is a struct member, the caller could modify elements in the slice, breaking struct invariants. 2. Returning a copy of the slice is safe, but adds a lot of overhead. What you'd actually want is to return an immutable slice, but Go does not provide any facilities to do so (apart from wrapping a slice, but the lack of generics and operator overloading makes this tedious).
I guess a lot of Go code will just assume that returned pointers/slices/maps will not be used in a way that breaks invariants. But you usually end up reading the source code of 3rd party packages to see what is safe, whereas in other languages you could just read the method signature.
tl;dr: I think ownership and preserving invariants usually give the most mental overhead and Go does zero in that department.
The funny thing about Go is that it makes up for its long list of weaknesses with essentially one single strength. You can actually read other people's code without much introduction to the concepts used in that codebase, because the number of possible meanings of any particular expression is much smaller than in other languages.
There is so much talk about Go being for dumb, second rate, corporate developers, because that's what Pike essentially said at one point (perhaps without thinking first).
But in fact, it's not the developers who are dumb. It's the process by which large corporations employ and dispose of developers. They are thrown into some project and expected to "hit the ground running". There's no time for explanation. So what they do is read code to acquaint themselves with the codebase and hopefully become productive before they move on to the next job. And that is the one task where Go really shines. Reading arbitrary pieces of code.
Of course powerful abstraction features eventually make reading code easier as well, but only after having learned the abstractions created for that particular problem and codebase and only if those abstractions are very carefully crafted.
Powerful language features help writers of code long before they help readers. And that, I believe, is essentially the dirty secret that Go exploits.
We all want to be brilliant writers of code when in fact we are often readers poking helplessly at half understood code to make something happen. We even forget our own abstractions once we haven't looked at them for a couple of months or even weeks.
The problem with Go is that it not only acknowledges this state of affairs, it also enshrines it.
Definitely. This really shines in the standard library, it consists of extremely readable code and is a good way to get up to speed on canonical Go.
We all want to be brilliant writers of code when in fact we are often readers poking helplessly at half understood code to make something happen. We even forget our own abstractions once we haven't looked at them for a couple of months or even weeks.
Definitely, but what are we comparing to? I would agree that e.g. C++ and Haskell have this property. Unless you understand the language and commonly-used abstractions, template-heavy C++ code is difficult to read. However, there are many languages that have more powerful type systems than Go, but where code is still easy to read (ML, Object Pascal, Oberon, Ada, etc.).
I wonder whether it is simply a theoretical tautology that the more abstraction features you have in a language, the more different possible meanings any particular syntactical expression can have, and the more effort it requires to figure out its true meaning, assuming you're not familiar with the codebase.
Or is that a false dichotomy? I am unfortunately not familiar with Ada or Oberon and Pascal is but a faint memory.
Sure, I could use something like Haskell, but then I'd have to worry about whether I'm accumulating a giant stack of thunks that'll blow up. go just works, and less of it will bit-rot than the comparable python script, thanks to at least some types.
This is why the SREs seem to be fans, e.g. https://talks.golang.org/2013/go-sreops.slide#1
I wrote some machine learning tools in Go and the experience is quite bad. The lack of operator overloading and parametric polymorphism make most ML code ugly. It also does not help that Go's compiler backend does not optimize very strongly and calling out to C comes with a relatively large overhead.
Sure, I could use something like Haskell, but then I'd have to worry about whether I'm accumulating a giant stack of thunks that'll blow up.
There are many languages between Go and Haskell that are productive and provide a sufficiently strong type system.
At a previous gig, I had done a lot of F#, and I think that's close to my personal sweet spot, but I'd have to use it frequently to keep it in my head. Go is small enough that I can load it into cache when I need it.
Definitely! I have continued to use Go for small utility every now and then as well. The standard library is extremely well-suited for that kind of work.
(Now using Rust more in that role as well, but mostly to get continued practice ;).)
Too many features.
Unless the project is very badly configured, that should be all you need to compile it. Now, writing those .gradle files...
Compiling most Java projects is usually as simple as 'brew install maven; mvn package'. Maven does require a build file, but it can also do more than the go utility, such as deploying artifacts in a repository, build distributable tarballs, RPMs, debs, etc.
Maven, like go, largely relies on convention over configuration as well. If you generate a POM file from the standard archetype, you basically drop your files in the right source directory and it will build.
(I am not a fan of Java, but I think a lot of the criticism here is lazy. The first time you use the 'go' tool you have to learn its usage as well: how to separate library code from programs, what are the conventions for package structure, how to avoid API breakage for downstream users, etc.)
Gradle build files have confusing syntax. Just knowing which lines have an equals sign and which don't is a bit of work.
Iterative languages seem to match more closely how people speak/think in verbal language.
I want more enforced required clarity. I compared it at the time to writing English without any punctuation, you can do it but it makes comprehension much more difficult.
The Go language doesn't have to be crappy to provide batteries included API, version stability, and great tooling. No generics and no sum types? Those are no longer groundbreaking, they're the bare minimum.
Old languages like C at least had the excuse of being created a long time ago, when we possibly didn't know better (and computers were much slower). Go's creators however don't have that excuse. They screwed up, plain and simple.
I love Python but always wished for a simple, type-safe language; Go gives me that. It's not worse than Python IMHO.
I only advocate Go as a replacement for those that would use C for user space applications, or possibly some kind of low level stuff.
If Python and Ruby eco-systems had blessed compilers, instead of just CPython and MRI, I doubt people would be flocking to Go.
We already see this happening in the Ruby world, just let Crystal become a bit more mature.
Btw I don't consider myself a Go developer, at work I use C# and Python. I would love to learn Rust, used to be the other way around, but I've lost hope in Go and have taken a second look at Rust, I love the direction Rust is going overall, but I understand why Go got so popular so quickly, it came at the right time with the right amount of working parts.
Just because a language is founded on good programming theory doesn't make it unviable in the industry.
Ideally, the reverse would be true.
Galois, of course, is very much built on Haskell. And there are certainly other examples.
A company doesn't need to be built on a single language.
I mean the fact that there is even an exhaustive "Haskell in industry" page at all shows you how rare it is. There isn't a "C++ in industry" page! I did actually find a Go one here:
https://github.com/golang/go/wiki/GoUsers
But it's both hilariously long and also obviously not exhaustive.
Maybe you should spend some time learning how FB and Microsoft use Haskell, you might be surprised with what you find.
As a clue, try to find out how C# and VB got LINQ, maybe you will find the right MSR paper by a certain Erik Meijer.
But as C++ programmers obviously never care about readability and simplicity, I am pretty sure they will take Rust. After all, Rust is a decent language and as users, we will all benefit from the migration.
Whereas with Go code, I have to filter out all the error handling (which often takes up 2/3rds of lines of code, even when it's just "if there's an error, return the error" which it almost always is), wade through the mass of functions that take interface{} and cast it to something internally, etc etc.
I've tried programming in Go and I find it horrifying - I either get to ignore errors or spend ~80% of my code doing this, repeatedly, a few times per function:
foo, err := bar()
if err {
return nil, err
}I know that error handling in Go can be tedious, but to some extent, it is also the programmer's job to utilize the features of the language to write clean code: https://blog.golang.org/errors-are-values
Regarding your type problem I find very few cases where I need an empty interface (mainly container types). Most of the time I use real interfaces and the most code I have seen used non-empty Interfaces too.
> Regarding your type problem I find very few cases where I need an empty interface (mainly container types).
I've run into the issue with quite a few "generic" libraries, since Go doesn't have a generics system. For example, https://github.com/manyminds/api2go requires a few of them in common cases.
My alternative is writing the same code over and over, and hoping that I've not missed something in one instance of it.
foo, err := bar()
if err { return nil, err }
But then the Go community yells at you for daring to not using `go fmt` style.More correctly, you would decorate the error:
foo, err := bar()
if err != nil {
return nil, fmt.Errorf("failed to do bar: %s", err)
} let foo = bar()?;
If you need to do something custom in converting bar()'s error into your error type - which is rare - you can use .map_err() like this: let foo = bar().map_err(|e| MyError::BarError(e, extraContext))?;The concept of not having features because someone on your team might use them? You subtly say that Rust is a worse language because it doesn't seem like it was designed in the 80's, when clearly as a language (together with the compiler), it's objectively better than Go.
So this is my personal preference and should not offend any Rust fans. Besides the syntax, which reminds me too much of C++, I think Rust is a very good language. I do not know if it is 'objectively better than Go.'.
Btw. restrictions are not always bad.
As a result, the "maturity" required for adoption differs in different parts of the C and C++ community. If Rust becomes more popular than either of those languages it is unlikely to be on a time scale shorter than a decade or two. That's not to say it can't happen quicker, especially given that communication and adoption can be much quicker today than it was even 15-20 years ago, but it's unlikely.
The Rust team showed great foresight for bringing about those changes before the initial 1.0 release and it will pay off in spades.
It's a key step in removing barriers to entry:
https://www.joelonsoftware.com/2000/06/03/strategy-letter-ii...
The other thing you want is proven, successful use in high-assurance systems. That is, systems that either didn't fail or provably couldn't in certain ways. These are almost all written in a subset of C or Ada/SPARK. The advantage of using those is you can combine them with a vast array of proprietary or open-source tooling to catch about any error you can think of if it's implementation. There's also formal specification and protocol analysis tools that combined with expert review can catch the rest. Rust, although a good choice for increased safety/security, doesn't have such tooling yet. That means they will get less correctness overall and over time vs MISRA-C or Ada/SPARK unless similar ecosystem in industry and CompSci emerges for Rust. That's why I recommend against it for high-assurance security for now.
It does seem good for medium-assurance security where you want to knock out low hanging fruit in systems code. It will avoid serious errors in C while providing additional benefits with type system and other features. Ada 2012 + SPARK 2014 are the standard for safe systems since they systematically eliminate all kinds of errors with a consistent design and tooling with decades of field success. I haven't seen a direct comparison with Rust on each protection to see if it matches it already or not. The main advantages Rust has over them are its borrow-checker for temporal safety, more usable method for safe concurrency, and (best for last) highly-active community to provide libraries or help. Go has similar benefits if its GC works with your use case but a lower, learning curve & possibly lower efficiency. Due to ecosystem benefits, these are main two I'm recommending for medium assurance if Ada/SPARK are too much to learn.
Looking at the MISRA-C guidelines (or more specifically, a pirated copy - are these available legitimately to the public?), it seems like about half of them aren't problems in Rust, because it's warning you about stupid things in C that can't be fixed for legacy reasons, and half of them are things that could be caught with a linter / static analysis tool like https://github.com/Manishearth/rust-clippy . Do you think building something that implements as many of the MISRA-C checks as possible is useful progress towards the high-assurance Rust goal?
We have to live in the world we have too (e.g. Signal on an Android is better than not having it, etc.). When I document side channel issues in customer code there is almost some way higher priority issue they have to fix first, etc. But your advice is good for someone endeavoring to do a "correct" approach from the ground up.
Thanks for the detailed thoughts though. This is the sort of thinking anyone who thinks "I will write a secure chat program" has to do ages before they start writing code.
Still need a language that allows such analysis at least for those components being analyzed.
Java is getting an AOT compiler in July (Graal, http://openjdk.java.net/jeps/295) that will let you AOT compile parts, or all, of your program, including the JVM modules themselves. This would seem to leave the GC as the main source of side channel vulnerabilities. The GC itself will become more pluggable as well, with a pure Java implementation. What requirements would you put on a GC for side channel safety? (Ignoring the obvious just allocate a large heap and never GC anything, eventually just restarting the process when it runs out of heap). What if we had a fully concurrent, pauseless GC (e.g. Azul), would that change things?
What other issues would remain in your opinion?
Before answering the other question, it helps to understand what covert channels are fundamentally. First, know there's always two parts: Sender w/ access to secrets but inability to do I/O outside the system; Recipient w/ no access to secrets but ability to send its info out. These might be in separate processes, partitions/VM's, or even on a network w/ stuff happening due to protocol interactions. Idea is to find a way to communicate that wasn't intended for communication. Should help confirm the dark things I say about mainstream INFOSEC when I say I couldn't find almost any intros or blog articles on these for you in top results. I found one, though, that describes them well even if not having many examples:
https://arxiv.org/pdf/1306.2252.pdf
Now, back to the GC. The GC might kick in whenever secrets are being processed. This could mask information about them due to unpredictability or leak information about them. The simplest route to dealing with it is not allowing GC while performing any operation that processes secrets. Depending on the app, the impact on memory availability or performance can vary considerably. There's also at least three more channels to look for on even basic app. The keys might leak in memory that's released back out of the system. First needing overwritten is in Java app itself for anything GC releases. Gotta look at assembly since compilers sometimes get rid of that as a "useless" operation. Second, the OS itself might leak by swapping out privileged Java app (Sender) for a Recipient that simply reads the registers before doing anything. Secrets might still be in them. Orange Book & separation kernels required these overwritten every time a process change. Don't know if current OS's do that. The "swap" part of the filesystem itself is a risk and should always be disabled. Finally, a Recipient in another process can receive secrets through the cache activity of the Java app. That's an old one that's hard enough to [confidently] deal with that I just advised running trusted apps on one CPU and untrusted on another CPU. That's physical isolation with them communicating over a pipeline, SMP since caches are separate, or multicore with no shared cache between cores. There's CompSci work such as partitioning caches and potential with embedded CPU's w/ things like locking, real-time caches that might help. Who knows real practicality, though.
So, you prevent or overwrite any storage of secrets that another process could touch. You put a brick wall between two events where a recipient observing timing could possibly learn something. Eliminating non-determinism can go a long way. Reference counting might help since you at least know when you'll deallocate w/ similar checks happening constantly. The deallocation might even be masked. For protocols, the classic response by military systems was fixed-size, fixed-rate transmission with extra attention that error responses didn't leak anything. Every detail you can sealed instead of just data payload since any might be storage channel. So, there's you a start on it.
And this was just running an app + considerations of a GC running in the background. Leads to all those problems. See why high-assurance security invested so much effort into automatically or at least reliably getting these damned things out of our systems? Just imagine how many are in UNIX API's, common protocols, and clouds. The fix is hard and expensive if it's legacy so they're definitely still leaky. :)
Or Sebastian was very dedicated to the joke, but they seem to have posted a comment confirming it's serious on /r/rust: https://www.reddit.com/r/rust/comments/62o9rx/tordev_tor_in_...
Looking at their comment history, both the Tor membership and the interest in Rust are there, and Sebastian just tagged themselves on the relevant Tor issue: https://trac.torproject.org/projects/tor/ticket/11331
But maybe this joke is only funny to those who enjoy programming in C? :)
It's just the ML recap mail which hit a rather poor timing.
Edit: cgo != Go. Thanks for the responses. I have done a bit of Go, but just pure Go.
A pure-Go rewrite might be an option (in fact Tor seems pretty firmly in Go's use cases), but that's not what the Tor team is trying to do.
[0] https://dave.cheney.net/2016/01/18/cgo-is-not-go
[1] a cgo->c call is ~100 times more expensive than a go->go call, and ~400 times more expensive than a c->c or rust->c call https://www.reddit.com/r/golang/comments/3oztwi/from_python_...
[2] https://www.cockroachlabs.com/blog/the-cost-and-complexity-o...
But yes, Go pointers in C code is bad juju.
Sorry, my reply was too brief. I wanted to add that the ergonomics are bad, not just because of freeing memory (for which the inconvenience can indeed be reduced with finalizers and/or Close() methods plus defer). But rules such as this one make Go->C FFI unergonomic as well. To give one example: many linear algebra libraries (e.g. Tensorflow) have their own wrappers around raw arrays to represent tensors (with their dimensionality) [1]. As a consequence of this rule, one cannot just a pointer to the first slice element to such functions (since a pointer to a Go object would be stored in a C struct), but have to malloc an array and copy over data from the slice to the C array.
[1] There are other issues, such aligning slice memory to 16-byte boundaries.
(Of course, I think the change led to a better overall API. Go just helped me get there in a circuitous way.)
It turns out (or so I hear) that Google statically links everything in production, and has been using C++ as a language to implement HTTP endpoints for a long time. So Go is a better C++ for what they want out of a better C++; for the rest of us, it looks more like a compiled language along the lines of Python/Ruby/etc. with a nice deployment story. If you want that out of your better C++, Go is great. If you want to reimplement all of Tor from scratch, Go certainly seems like a reasonable choice.
But as a result of these priorities, Go basically doesn't have interoperability with the platform ABI as a goal. (For some combination of historical reasons and the lack of complicated features in C, the platform ABI on just about every platform these days is a C ABI.) Rust does; it uses a standard compiler toolchain (LLVM) instead of what's basically a custom one (Plan 9), and the standard toolchain knows how to generate calls that follow the C ABI. Rust doesn't have a runtime of its own, and it's safe to directly call into a Rust program from some arbitrary point in a C program. Rust's allocator doesn't care if you do stupid things with pointers it allocates, as long as you give them back eventually. Rust doesn't create threads on its own unless you ask. Rust functions use the normal stack. Rust on UNIX uses the platform libc. And so forth.
It's possible to call C code from Go and vice versa, just as it's possible to call C code from Python and vice versa. But Go is not best tool for this particular job.
+1 for the rest of the explanation.
Rust has some interesting features, but being C isn't one of them.
What would be the advantage, just improved memory safety guarantees for people working on the project?
If that's the case start again from scratch and the first thing you do is build a handle based memory management system that everything must go through Safexxxx() versions of everything and then explicitly enforce no other memory access patterns.
If you want to do concurrency, build simple threads with input queues that would just be like channels in Go.
You can use a completely functional actor model in C. Is all this object oriented dangling pointer stuff that is scary, but there really isn't any need for that.
Software pipelines with slab allocation or ring buffers virtually guarantee no leaks and are very scalable and cache friendly.
That would be my personal recommendation.
The advantage of Rust is, honestly, that it has a community of people who are excited to do this sort of work in a language. At least 50% of the advantage of any particular language is not the benefits of the language directly, but the community around it (that derives indirectly from the benefits of the language, but also from marketing and other things). Some time ago I needed to write bindings to SCM_RIGHTS / CMSG_*, which is probably the trickiest part of the libc API (all of cmsg(3) is macros that do stupid things with casts). It was annoying, but it led to a better interface than C itself offers, and someone else found a bug in the OS X handling. I couldn't expect that if I were working on my own custom reimplementation of libc.
And is funny you talk about community, since as the guys on slashdot and others pointed out, the community of people who can use and work in C effectively is orders of magnitudes more than Rust. Have you any idea how many Linux Kernel developers there are alone?
http://m.slashdot.org/story/324469
Building API's specific to the use case is part of our jobs as professional developers.
Abstracting a few things things isn't "rewriting libc", that is just general practice for most decent size projects.
Anyway, the decision is made, so the whole thing is moot at this point.
The key word here is "effectively." For the purposes at hand, effectiveness includes memory-safety. Do you know of a single project that implements cmsg(3) in C or C++ in a memory-safe, well-typed, cross-platform way? Or a project where I can submit a pull request and expect it to be reviewed, tested across platforms, and fixed?
I do genuinely believe that the community of people who can use and work in C effectively, in the sense of effectiveness that I and the Tor Project are interested in, is orders of magnitudes smaller than Rust.
I have a very good idea of how many Linux kernel developers there are - and also how many high-severity security bugs there are. I'm a coauthor of a research paper where we wanted to talk about exploitable security bugs in Linux, so we sat down and found a local privilege escalation in hours (CVE-2009-0024).
Rust provides, among other things, a decent type system that can encode some of your invariants in a manner that's a lot clearer - and every operation goes through the equivalent of your Safexxxx() functions by default, making the actual logic of what you're trying to do further clearer. A reasonable error handling mechanism is also built into the standard library and partially into the language, rather than checking return values for -1, or is it 0, and is the actual error in errno or was that function actually meant to return a value between -1 and -100 corresponding to the error, etc etc.
Therefore a slow transition of rewriting parts of the code in a safer language and having the core still in C is much less feasible with Go. With Rust you can easier just compile some object files and link them into your application.
http://www.adacore.com/uploads_gems/07_safe_secure_ada_2005_...
https://en.wikibooks.org/wiki/Ada_Programming/Tasking
http://courses.cs.vt.edu/cs5204/sp99/Overheads/6UP/6UPCSPand...
So i read documentation instead. Ada mostly seems like a pretty sensible language. Its story on memory safety, though, seems to be "don't use dynamically allocated memory", or at least, if you do, you're on your own. It doesn't have anything like Rust's safety guarantees, or those that fall out of having a garbage collector.
If it was 2007, and Ada was a bit more easily available, I'd be agitating for it. But i think its shot at the open source big time has passed.
http://libre.adacore.com/download/
You can download the 'GPL edition' of GNAT [1]. This is a complete toolchain for Ada 2005, with the same actual compiler as the commercial version, i think. The libraries are GPL'd, so if you distribute a binary built with it, it has to be GPL'd. So, no Ada 2012, and no way to distribute binaries willy-nilly, but certainly enough to explore the language.
I always puzzled why even C++ does not provide any facilities like that out of the box.
(2) Compiler and runtime licenses. I always feel like the version of Ada I get is either somehow not the "best" one or has some lingering license issue with the runtime library. I don't think this is necessarily true, but it's easy to get that impression.
(3) Packaging, libraries, and documentation. Many developers expect a tool like cargo to be available and have access to a wide variety of open-source libraries. They also expect lots of tutorial documentation and blogs.
"Memory Management with Ada 2012"
2) Unix, C is so fundamental to building software, that I think any language that doesn't share syntax with it is doomed. Having a common syntax helps in learning new languages, IMO, and can also be a launching point for differing semantics...
EDIT: I keep thinking a different language that acts as a front end w/ a better syntax might be a good idea. It outputs Ada that integrates with the tooling ecosystem. Also, a seemless FFI for C libraries like Julia's.
It certainly wouldn't hurt if people simply used a more literate programming style no matter what language they choose
Far too many people think code as below is acceptable.
This is C obviously, but pretty much equal horrors around in every language.
This isn't 1994 and the compiler really doesn't care how long your variable names are, plus EatWhite() is pretty damn fast.
const int to_pn = base_n ^ label_n;
const int from_p = _array[to_pn].check;
const int base_p = _array[from_p].base ();
const bool flag
= _consult (base_n, base_p, _ninfo[from_n].child, _ninfo[from_p].child);
uchar child[256];
uchar* const first = &child[0];
uchar* const last = flag ? _set_child (first, base_n, _ninfo[from_n].child, label_n)
: _set_child (first, base_p, _ninfoNo, that's C++.
struct {
int check;
int (*base)(void);
};You mean Rust I think ;)
So, Rust ain't the Ada makeover Im thinking about. It's also in a stability-oriented freeze of existing design right now. So, the makeover will need to be a different language.
I won't lie to you. It took me longer to produce good code in Rust than almost any other language I've recently learned. But it's worth it; it's quite simply amazing that it is capable of guaranteeing what it does, with nearly zero overhead. Also, I'd say in many ways it's simpler than C, because it has no undefined behavior; I don't need to concern myself with any of the nuances that I had to learn in C. I highly encourage you to give it two weeks, that's what it took me to become hooked.
There's also a penalty a language pays for following a C-like syntax and semantics but diverging in specific but significant ways. In this category, I present Perl, which much of whole generation of users decided to treat like C, and got very confused and upset when it didn't always behave like they expected (which I maintain is because they didn't actually understand the language as well as they thought). There's a penalty for being very like something else to the point that people can mostly ignore the differences, but occasionally those differences come out to bite them if they haven't actually learned what they are.
Booking.com's soup du jour if you will.
Basically, if you aren't using JavaScript in a web shop, and you claim to be a "full-stack" "ninja", then probably you are using Go around here as a jobbing programmer.
I know that sounds terribly cynical and obviously a massive generalization but that is my personal experience.
I don't have a problem with Go per-se, but I do have a problem that a lot of people seem to go to extreme lengths to defend what someone else mentioned is frankly a pretty "mundane" language, citing memory safety, but more strangely portability and performance as its wonderful virtues.
When I meet with the zealots of the Go community around here I often have to quickly excuse myself with good grace.
Seriously, if that's what you want... go (pardon the pun) use C#, it is a saner and more expressive language, has a higher performance runtime and is way more portable with less vendor / career lock-in.
Go... I just don't get it.
Positives... Has a fairly nice package manager like npm.
Isn't made by Microsoft if that's your thing.
Oh and heaven forbid you don't have to think much... until you do because it is slow.
For a project like Tor though, which is damn slow as it is, I think it is totally the wrong choice.
Developers, please suck it up, get over it and learn C/C++ and some variation of Lisp.
Fine, go play with different languages for fun stuff. Use python for ML, try serious meta-programming in D, or jump into Haskell for kicks.
On the other hand modern C++ really can be a very safe language to work in if you can be bothered, and you are only deceiving yourself and your project going with something less sympathetic to the machine itself, especially if you are working on infrastructure level systems.
Sorry if this offends anyone, and of course it is just one opinion, but I think I'm being fairly nice as compared to what Linus might have said in comparison.
Basically stop it with the hobbyist shit. You are horrifying me and probably many others.
You aren't writing a web page here, so please treat the project seriously.
IMHO
I've got over 20 years of writing production C/C++ under my belt and I know Lisp. So I've "sucked it up." Am I allowed to like Go now?
It never ceases to amaze me how many people are bothered about other people's taste in something so mundane. If you don't enjoy programming in Go, don't do it. I personally think it feels light and easy like a scripting language, but with more C-like performance.
I like the fact that its stdlib is very complete, so I can sit down and do pretty much any kind of small project with zero external dependencies, I like the fact that the core libraries are well-designed so the interfaces are consistent and easy to learn, I like the fact that it produces statically-linked binaries so deployment issues are minimal, I like the simplicity of the CSP approach for a lot of concurrency problems. I could also come up with a list of things I don't like about Go, but I'm not going to bother, because I've decided that on the whole, I really like it as a tool.
Btw: > Positives... Has a fairly nice package manager like npm
One of my biggest complaints about golang is that it really doesn't have nice dependency management. It's terrible and probably the one issue that has made me seriously consider walking away. :)
Ordinarily, I would agree; I don't really care what people do when writing application code. The way it's filtered into devops tooling makes the choices of Go peoplea problem for me, though. If I'm going to be stuck with a language with bad error handling and inexpressive typing, I'd rather it be Python or Ruby so at least I can leverage dynamic typing instead of bad static typing.
You are right, the standard library is a pleasure to work with in Golang and for the most part I trust it, but hey anything you can do in that library is easily reproducible in more or less any other mainstream language.
I really don't have an axe to grind with the language itself, it's a fun and productive thing to use for sure.
My main problem is with people who spend most of their time writing customer service applications thinking that the same language they are ultra successful with in that space is suddenly appropriate for writing systems software, because, ughh "safety"?
There is a heck of a lot more to safety than buffer overruns, and if you can't formalize abstractions for dealing with memory usage patterns, one might argue you have no business writing Tor in the first place.
Hah well compared to C/C++ dependency management, Golang does pretty darn well!
http://benchmarksgame.alioth.debian.org/u64q/compare.php?lan...
[1] https://github.com/golang/go/wiki/PackageManagementTools
http://www.cvedetails.com/product/5516/TOR-TOR.html?vendor_i...
This list is full of memory safety issues.
_exploited_ is the operative word here.
Show me the list where people wrote exploits for the bugs you point out, or where someone abused them to de-anonymize a Tor user? There aren't any.
[1] shameless plug: https://github.com/duneroadrunner/SaferCPlusPlus
Second, many websites (sadly) do not work if you are using a VPN (like Netflix).
I guess if you're open-source and a lot of people are paying attention, that would do it.
Also, what about the Netflix issue?
I really just want things that I access in Private Browsing, or using a service other than http/https, not to be recorded.
I understand that people in various countries want to use the Netflix from another country. But I'm representing the 99%.