Space Monkey dumps Python for Go
spacemonkey.com
spacemonkey.com
I've been using Go for almost two years now. Like OP, I am/was primarily a Python developer. Like OP, my first Go project was a time-sensitive rewrite[0] of a project from Python (tornado) to Go.
Even though I was an experienced Python developer, the Go rewrite was marginally (~20%) faster[1]. But the real benefit came from the subsequent development - refactoring, rearchitecting, and maintaining the project on an ongoing basis. Go was designed to make it easy to scale the maintenance[2] of software, and on this one axis, it absolutely blows every other language and environment I've used out of the water.
For a fresh project, I'd say Go is about 10% slower to write than the equivalent Python[3] for the average case. But the time/cost savings are very quick to come thereafter.
[0] I would absolutely not recommend doing time-sensitive rewrites in general, but that decision was a separate matter.
[1] Some of this is due to the nature of rewrites in general, but the fact that it wasn't slower to use a language I'd never used before says something about the language.
[2] Scaling software development as teams grow is very different from scaling software as users grow.
[3] Assuming comparable familiarity with both languages, which is rarely the case, of course
(disclosure: I haven't fiddled with Go at all yet, so I know nothing about the language itself)
- super fast compile times for fast developer iterations
- you can create an interface (set of methods) that your module owns, and apply it to objects created by other modules without recompiling. If you only need one method you can take that one. It encourages decoupling you from your dependencies.
- the code has one correct formatting convention and a tool that will auto-format your code
- many complex rules for numeric conversions that are implicit in C and cause no end of trouble are explicit and much simpler in Go.
- treating concurrency as a series of sequential process connected with channels makes it MUCH easier to reason about.
Beginners usually think Python strong indentation is a weakness of the language. I actually find C++ freedom being more error prone:
if (a < b);
a = b
It's a not very frequent bug, but when it happens it takes you hours to spot. ;) if (a < b) {
a = b;
}EDIT: Whoops, it's actually compiled into bytecode then executed by the VM.
https://en.wikipedia.org/wiki/Python_(programming_language)#...
A quick Google search brought up this, for instance:https://bitbucket.org/StephaneBunel/pythonpep8autoformat
Yes we can also make a single big ass executable out of any python code but it's slower and not supported out of the box.
One thing I really missed when I abandoned the Java world and went to C++/Python was the refactoring and formatting tools available in the IDE. I could set it up so it'd auto-format my code on save, so that I could rename methods with a single keystroke, so that I could pull out classes or add getters/setters, so that I could add parameters, etc. I know Go supports some of that through gofmt and go fix, but I'm curious if anyone is both an experienced Java dev and an experienced Go dev and can compare the continued maintenance costs of the two.
BTW, I would put the productivity premium of Go over Python at about 50%-2x, not 10%. I did a mid-size (~6 man-months, though much of that time went into interactions with other teams) green-field prototype in Go at Google, and found myself really missing constructs like list comprehensions, dynamic typing, strong support for literal data, ability to treat user-defined types just like the built-ins, etc. I do think you'll make some of that back in continued maintenance - I actually really enjoyed working with Go - but I didn't experience the claims of "Go is just as good as Python for prototyping", as someone with a fair amount of Python experience.
Gofmt is actually one of my favorite parts about go and I feel that it, along with the simple documentation system and godoc.org save me a ton of time making my code ready for others to read/use.
The code base behind gofmt and gofix can also be used for other tools, i haven't felt the need for a powerful refactoring tool but I'm sure one will show up eventually.
I do miss list comprehensions and operator overloading but I understand the dev's arguments against those.
Of those, I only see comprehensions as something that's severely lacking, though I am not a Python guru by any measure. Dynamic typing: Go has duck typing, which gives some of the same benefits. Different refactoring strategies that take advantage of duck typing over other things you'd do with dynamic languages (like just substituting an object of a different type willy-nilly) will probably make up for the rest. I don't see strong support for literal data as a big productivity gain. Go also does a rather good job of treating user defined types like the built-ins and vice-versa, with a few exceptions like maps.
I didn't experience the claims of "Go is just as good as Python for prototyping", as someone with a fair amount of Python experience.
A lot of your competence with Python for prototyping probably has to do with your competence with Python. A lot of the advantage of Go is what it doesn't have. It's a much simpler system, so it's easier to design around, especially as one needs to get closer to the metal.
Python just hands you back a dict of lists, and you pull out the information you want, no type declarations needed.
The "user-defined types like built-ins" lets you build libraries where you can use the basic language structures on the data structure itself, without having to convert to an intermediate form. For example, the Python bindings to Amazon S3 let you access files through dictionary notation. BeautifulSoup lets you iterate over child nodes with a for loop, or access attributes as a dict. NumPy lets you use standard arithmetic operators on matrices. Judicious use of this in libraries makes the resulting code much briefer.
> I was decoding a large JSON data structure that could
> easily have thousands of different fields nested 10+
> levels deep. Go's standard json library usually
> recommends that you define a struct type to hold the
> result of the decode, tagging field names with the
> `json` annotation. This obviously doesn't work when you
> only care about a few dozen of the several thousand
> possible fields.
Sure it does! type Response struct {
LevelOne map[string]struct{
LevelTwo map[int]struct{
Value string `json:"eventual_value"`
} `json:"level_two"`
} `json:"level_one"`
}
The response can have way more fields than just the ones you declare in your struct. response['level_one']['level_two']['eventual_value']
?I did ignore fields that I wasn't interested in; however, you do at least have to declare all the intermediate structs in the tree. With 10-ish levels of nesting, that's a lot of boilerplate; even in your example with 3 levels, you've still got 6 lines of code for what's a one-liner in Python.
It wouldn't be as much of a problem for a production system where you define your data structures once and then expect to amortize them over many, many changes. But I was prototyping, and the point of prototyping is to not spend time on boilerplate that isn't necessary to prove out your concept. A few hundred lines of type declarations just so I could decode JSON isn't exactly convenient.
But, at your insistence, we can keep digging. It's easy to write a polymorphic function that arbitrarily picks out values from a Go `interface{}` value.
func lookup(v interface{}, keys ...string) interface{} {
switch len(keys) {
case 0: return v
default:
dict, ok := v.(map[string]interface{})
if !ok {
panic(fmt.Sprintf("%T is not a dictionary", v))
}
return lookup(dict[keys[0]], keys[1:]...)
}
}
And then we can use it like you would in Python: lookup(response, "level_one", "level_two", "eventual_value")
This doesn't completely remove the burden of type asserting (unfortunate), but it at least makes nested lookups much easier. Arguably, the nested lookups are probably where the type asserts hurt the most.Full example: http://play.golang.org/p/4tWtqGRgU6
Yeah, prototyping with JSON is one place where dynamic languages are more nimble.
For example, the Python bindings to Amazon S3 let you access files through dictionary notation. BeautifulSoup lets you iterate over child nodes with a for loop, or access attributes as a dict. NumPy lets you use standard arithmetic operators on matrices. Judicious use of this in libraries makes the resulting code much briefer.
Yes, this is another specific area where dynamic languages do prototyping better and can produce shorter code. I did an exercise where someone implemented a routine in Clojure that I had written in Python. The Python was shorter! In part, this was due to list comprehensions. In part, it was due to an excellent library. What you describe is what many Smalltalkers did as well: create your own control structures and DSLs.
I've sometimes wondered what could be done with a dynamic superset language of Go, such that one could mostly add type definitions and compile to Go.
It should be simple enough to implement a dynamic JSON parsing DSL in Go. It would be slower and memory inefficient, but much more convenient for rapid development.
Go doesn't have a REPL and that is objectively a serious dent on effective prototyping.
On formatting, gofmt is awesome. There's an IDE called LiteIDE which automatically calls gofmt for you, or you can just run it yourself whenever. Additionally, because there's deliberately no room to customize gofmt, you don't have to spend time in pointless meetings when "that guy" decides it's time to rewrite your company eclipse templates and argue over all the minutiae.
As far as refactoring, nobody's built an IDE integration that comes close to eclipse/IDEA for that as far as I know. It's totally doable in Go, in a way that it wouldn't be for Python (yay types), but I don't think anyone's done it. On the upside, doing it manually is less painful than it would be in typical Java code, due to type inference and having less code to begin with.
The reformatting is actually quite aggressive (it will add/remove blanks and move braces and bits of code around), and so does a good job of keeping everybody's code looking consistent. The Visual Studio text editing functionality only occasionally reaches even the lofty heights of "pretty good", but I do think the document reformatting actually even exceeds that.
(Shame the settings are are per-user rather than per-project, but there you go. It wouldn't be Visual Studio if they didn't manage to fuck something up at the last moment.)
For example,
[1] https://github.com/DisposaBoy/GoSublime
or on emacs, if you use the packaged go-mode.el, you can run gofmt on the "before-save-hook" to do that for you...
It's slightly annoying to setup, but absolutely worth it.
http://blog.repustate.com/migrating-code-from-python-to-gola...
...We spent long nights poring over profiling readouts and traces, wrote our own monitoring system, and wrote our own benchmarking tools
... We enabled cgroups on our Python process to tightly control memory usage and swap contention without requiring a different Python memory allocator scheme.
This is very interesting for the rest of the Python/Twisted community - now that they are not using this (and so is'nt part of the secret sauce), I wonder if they are willing to push this out as a patch/blog post ?
really look forward to this..
Any idea where the overhead comes from? Twisted, or the Python interpreter itself? Is this a GIL performance issue? Or perhaps even lower -- something here is really hostile to CPU cache?
I realize this is a matter of taste, but my favorite async framework is still the kernel. Write small programs that do one thing well (and thus have pretty uniform workloads) and then let the kernel balance resource usage.
After continuing to hit walls with the standard Python Protocol Buffer library, we wrote our own that was 5x faster than the barely-documented C++ Python Protocol Buffer compiled module support.
Ugh yeah, the standard Python protobuf library is pure python and horribly slow. And it requires code generation -- in a dynamic language! The C++ one is faster, but also requires code generation and is just nasty to work with.
Not that this matters much to you at this point, but I have a small C/Python protobuf library that's 10x faster than the standard Python protobuf: https://github.com/acg/lwpb
PS. I see you're in SLC area -- me too. We should talk tech shop in person sometime!
I think the sort of performance issues we were hitting were honestly related to just Python runtime overhead. A Python function call is actually really expensive, and with Twisted Deferred handling, it's really hard to eliminate the massive amount of function calls each I/O event does.
We just had a really slow CPU so we did our best to eliminate as many Python function calls in hotspots as possible, but yeah that was challenging.
Ah, makes total sense. Any particular reason you chose to use Twisted in a hardware appliance? Is there a web interface that's supposed to be super responsive?
You can find contact information at http://baroquesoftware.com/
Hindsight is 20/20 I guess.
I have my doubts that that would have solved all of our problems though, but that's certainly a different path we could have taken.
I recognize that Golang is filling a gap, letting people get static typing without having to learn much new stuff, and there's value in that, but I think Go enthusiasts should read these great posts by Tikhon Jelvis with an open mind:
https://www.quora.com/Googles-programming-language-Go-seems-...
https://www.quora.com/Go-programming-language/Do-you-feel-th...
This isn't a revolutionary thought. More compile time safety generally means more complexity in the type system. We, as programmers, must decide when, where and how much we're willing to pay for that type safety.
You can go balls-to-the-wall with Idris or Agda or ATS and encode a whole mess of things directly into your types. Maybe you can completely remove out-of-bounds errors! Surely, this is better than the alternative, which admits less safety.
This isn't a fair comparison because all three of those languages are ridiculously hard to use for practical things. (Well, it's hard for me, anyway.)
Then there are other, more practical languages like Rust, Java, Haskell, OCaml, etc., that all provide a helluva lot more compile time type safety than Go does. (Or, to be precise, a lot more expressive power with respect to compile time safe polymorphism.) Some of the languages go even further; Rust eliminates data races and Haskell isolates "side effects" with monads.
All of these things are wonderful, because we get a lot more guarantees about ways in which our program cannot fail once it's compiled. But they aren't free. They come with a cost. Typically the cost is in greater language complexity. The programmer must write code that the compiler accepts. If you're in that language's wheelhouse, maybe the complexity isn't so bad. But when you want to write a program that isn't easily accepted by the compiler, you run into complexity. Maybe it's not that bad---maybe you just need to figure out how to write a monad transformer. Or maybe you just need to think a bit more carefully about the lifetime of a borrowed pointer. But other times, you need type families and multi-parameter type classes before you can make an unboxed vector with your own type: http://hackage.haskell.org/package/vector-0.10.9.1/docs/Data... --- The safety is great, but dammit, you're going to pay for it.
On some days and for some projects, I'm more than happy to pay the price for safety. It is immensely satisfying to write a program and be able to make all sorts of claims about its safety. (I just wrote a regexp implementation in Rust. Guaranteed no memory errors. No GC. Totally type safe. It's awesome.) But other times, I don't feel like paying the cost.
And that's OK. It's not sad. It's not "ugly." It's a legitimate and reasonable choice given a trade off.
Saying a language is "badly designed" just because it didn't include your preferred set of features is bad juju IMO.
It's absolutely true that there is always a trade-off. Metaprogramming makes code more difficult to understand and hence more prone to certain types of errors. That's true for runtime metaprogramming as well though.
So if the lack of generics leads to an overuse of reflection (or worse to custom code generators) then the complexity you avoided in the language is re-introduced through the backdoor.
At the end of the day I think it's about the culture that gets established around a language. Go's creators are trying really hard to socially engineer a culture of simplicity, and that's laudable. But it reminds me of the early days of Java even though the Go community would rather have me think of C.
The big difference: C had compile time metaprogramming from the beginning, Java did not. Look which one turned out to be the simpler ecosystem in the end.
Nit: they're typesafe (when done using interfaces), just not statically so.
If Python is fast enough for you, it’s a fantastic language. The problem is once performance or codebase demands scale, dynamic typing rears its ugly head and there are no simple solutions. At work we sidestep this issue by writing a plethora of tests, but now dynamic typing productivity gains are offset and we spin up a lot of AWS instances for performance.
[0] Go, Rust, Scala. Haskell and OCaml have had it for a while.
I guess the concept here is that Go is between Python and Java: (much of) the ease of use of Python + (much of) the performance of Java. The type system literally straddles the dynamic and static.
[ I also detect an enthusiasm for cool, new tech for its own sake (despite risks and cost, or whether it's actually better or not - hell let's find out!) ]
> The type system literally straddles the dynamic and
> static
The Go type system is fully static...No, that gives you a static type system with loopholes, same as Java, you have to use reflect to get a dynamic system in which you can invoke an arbitrary and unchecked method on an object (also same as Java).
Granted, because Go's type system is structural you could define an interface with just the method(s) you want, cast your object to that interface and then call the methods, and you can even define the interface anonymously and inline. Still, go requires your static types to match unless you use reflect.
Why? Because they did not choose Java?
Opinions vary highly on that claim. In fact, I've never seen anyone claim "ease-of-use" as one of Java's properties.
It took us about 4 weeks."
How many people worked on the project? What's the build time on a program of this size?
We'll be releasing a bunch of these. Particularly the OpenSSL bindings, which even without hardware acceleration are faster than Go's TLS library, and despite Heartbleed, OpenSSL is not vulnerable to timing attacks like crypto/tls is.
Without tests but with our supporting libraries, our codebase was 36784 lines of Python. Those same lines became 41717 lines of Go.
Counting lines of code is hard.
Sure, python has list comprehensions, but I can tell you how often I've seen hugely over-complicated one line list comprehensions that were impossible to understand unless you were absolutely sure what it was supposed to be doing before you read the code. I've refactored some of those in my lifetime just to make them more readable. That's not to say that most simple list comprehensions aren't totally fine.
So, just saying fewer lines of code makes for fewer bugs is not 100% accurate. I'd say, less /functionality/ leads to fewer bugs. But except in extreme circumstances (like writing a whole framework to do one small task), actual lines of code does not correlate to the number of bugs.
Our solution was to switch to a dedicated ARM board (beaglebone) attached to the router. But I'm definitely going to take a look at using a compiled language now, as the codebase is still very small.
Does anyone else think that this rings true? And if so - is it a cultural thing?
My only data-point is the shop I work for, and folks in the community that I know come from a Ruby background, such as Dominik Honnef, Ben Johnson, Mitchell Hashimoto.
Ouch. That was a lot of work for 0.2 MB/s. The next 2.8 MB/s was also a lot of work, but it seems conceptually more straightforward.
Python is a joy when it's fast enough, but it's not surprising that it isn't good at throughput on small devices.
I was surprised recently: first that Go itself bundles a cross compiler (you can easily download Go on x86 and compile ARM binaries with it), second that this made developing for Raspberry Pi and similar platforms very productive, and third that many Linux distributions have QEMU and set up binfmt translation so ARM-compiled binaries run on x86 with no extra effort.
For the uninitiated: https://wiki.debian.org/EmDebian/CrossDebootstrap#QEMU.2Fdeb...
here's the current bleeding edge: https://code.google.com/p/go/source/browse
here's the issue tracker: https://code.google.com/p/go/issues/list
looks like it's being used by the dev team.
If you would like to fork the golang project, run this command:
hg clone https://code.google.com/p/go
This is pretty clearly not Android.I never tried it, this would have been a nice use case.
A question here: Is it always a practical approach to transliterate line by line when transitioning to another language or because "Go [is] semantically very similar to Python"?
Much easier to convert first (with whatever warts the existing code base has) then do a refactor.
So I doubt their approach had too much to do with Go per se and more with wanting to get something out sooner rather than later.
We didn't end up using channels too much. Deferreds got translated into Futures that we wrote a small library for. Many of our Go-specific utility classes do use channels heavily though.
I suspect we can improve go runtime scheduler performance by starting to go through and replace pieces with channels where appropriate.
Internally, Google has something faster for Python, but they haven't released it yet. You might try putting pressure on them here: https://code.google.com/p/protobuf/issues/detail?id=434
http://yz.mit.edu/wp/fast-native-c-protocol-buffers-from-pyt... talks about the issues some more with some links to other rewrite attempts.
That said, if I did want to go looking elsewhere, I think the main serialization format I am really interested in is Cap'n Proto: http://kentonv.github.io/capnproto/
If you want to write code in a Domain Specific Language (even if that domain is "math"), it's nearly impossible in Go. You can't do operator overloading or all the really crazy hacking on the functionality of basic parts of the language to make it work in fundamentally different ways (the way you can with numpy etc). Go will always look like types, functions, and methods.
Right now, Go has no user-defined generics.. this is actually not a problem for many projects, but if you happen to need a lot of specialized containers for different types (like red black trees, graphs, etc), then it can be a barrier. There are ways around it, probably the best way is just code generation to duplicate the container code per type that you need. The other option is just using interface{}, which is basically void* or "Object" for Go, and casting in and out of it... but that can be slow.
There's probably some stuff I'm missing, but those are the major points that I can think of.
As a fan of DSLs, I actually do appreciate how easy it is to understand any Go that anyone has written, since there aren't any cutesy DSL tricks anywhere. So, yeah, it's hard/impossible to do DSLs, but that might not be a Bad Thing.
Generics is a (frequently discussed) downside. It's definitely a tradeoff. For every container type we had, we had to make a copy of it or a typesafe wrapper around it for every possible instantiation of the container with different types. If you're used to coding by making flexible containers and so on, that's definitely much more brittle and challenging in Go. Go peeps will tell you that's part of a tradeoff they're willing to take.
No offense, but that was a hideous amount of effort for just 20% in one area. An extra dollar or two spent on a hardware upgrade wins for everything.
What does that even mean?
If your code takes time proportional to 2^2^n then for n=4 that's 2^16 x constant, which is probably fine for any reasonable constant, and for n=6 that's 2^64 x constant, which is probably too much for any reasonable constant, so if you're running into problems with n>=6 then moving to a different language or faster hardware or more-micro-optimized code is not going to help you.
On the other hand, if your code takes time of order n^2 and typical values of n are in the thousands or millions, then the difference between C and Python, or Core i7 and Z80, or sloppily written code and highly-bummed code with an inner loop in hand-tuned assembler, may make a big difference.
So, having established that you were merely complaining about a grammatical slip, and that despite asking "what does ... even mean?" you knew perfectly well what it meant:
Yeah, I agree, it would have been more correct had they written something like "provided that your problems are not a matter of algorithmic complexity" or "provided that your problem is not algorithmic complexity". Congratulations, you found a mistake. (Though not, I think, anything to do with thinking "algorithmic complexity" an adjective.) But (1) so what? and (2) if I'm right in detecting a subtext of "man, if they can make that mistake, why should I take anything else they say seriously?", then I strongly disagree; anyone can perpetrate the occasional grammatical error. (Just out of curiosity, I looked at some of your recent comments, and found one having looked through considerably less text than the length of the article in which you found an error. I dare say one could do the same for mine. To err is human.)
Now that can't be a bad thing! SpaceMonkey needs to use it as a selling point.
Further, we have always really been inspired by Joel Spolsky's article on rewrites: http://www.joelonsoftware.com/articles/fog0000000069.html In fact, I can't help but wonder if we would have attempted a Go rewrite sooner if not for that article.
Perhaps we've come a pretty long ways in the interpreted languages department? Or, that the cost of interpreting is vastly less than speed losses due to IO or memory access?
"Interpreted is always slower than compiled lol right guise?" was a tired refrain ten years ago, much less today.
Maybe they'll make it faster in Python 3000, right?
Since it's a reference implementation it should be easy read and learn from.
They're not really willing to accept patches that speed up some things if they at the same time slow down other things (this has been the main problem with all the GIL removal patches that have shown up over the years).
They're not willing to accept patches that break any existing code or libraries.
PyPy on the other hand have non of these limitations and happily break all three, making it great for a subset of python code out there.
> (that explains why node.js is a bit faster than Go for handling HTTP requests at the moment)
Pretty sure node's http parsing is handled by http-parser[1], which is written in C (originally pulled from nginx as I recall).Go's http parser is written in Go.
[1]: https://github.com/joyent/http-parser/blob/master/LICENSE-MI...
> (that explains why node.js is a bit faster than Go for handling HTTP requests at the moment)
I'm not sure that is true any more:http://www.techempower.com/benchmarks/#section=data-r8&hw=i7...
At any rate, "handling HTTP requests" can't possibly be your bottleneck, given network communication, serialization, etc.
While it may be obvious, the interesting question is why exactly. There's been a good presentation from Alex Gaynor [0] on this topic. TL;DR: Dynamic languages ar slow not because they can't be optimized, but because developers make a worse use of memory and have sucky algorithms, inducing way more mem allocation and copy than needed. If you can be smart about those, then your dynamic language can match static languages.
[0] https://speakerdeck.com/alex/why-python-ruby-and-javascript-...
Using any of the Ruby implementations that don't have a GIL. Notably, JRuby, MacRuby/RubyMotion, and (I think) Rubinius.
This is not merely nitpicking. Any turing-complete language can be interpreted or compiled. Language specs tend to make one side easier, but that doesn't mean that Python cannot be compiled (and it is actually compiled to bytecode). You can have something similar to a GIL in a compiled language too.
One of the things that makes it easier to increase performance is having type declarations. This makes it much easier for the compiler to reason about the code, which can lead to increased efficiency. But optional type annotations can accomplish just as much.