How We Moved Our API From Ruby to Go
blog.parse.com
blog.parse.com
But I wouldn't recommend it to beginners trying to build a secure web application, because they will have to roll their own code when it comes user registration, authentication, authorization management, and i'm not even talking about writing thread safe libraries where one has to deal with mutexes and shared memory access.
If you need something highly concurrent and you know what you're doing, by all means. But Go is definitely not RAD for classic website development, and it's easy to shoot oneself in the foot with Go. It's by no mean a silver bullet.
I'm surprises Parse is coded in Rails,I assumed to was built on Nodejs, since it uses JS for cloud code.
I love Ruby and have worked in a high performance Ruby environment, but it can be pretty problematic to scale (and frankly, it was Apache and Solr that made the environment high performance, not Ruby). Clearly Go can scale, but would it significantly slow down early development?
At some point we're going to face this issue, and I'm sure others will too. I'd hate to reinvent the wheel for parsing double encoding etc. that you've already done.
We probably could have made a great jruby memcached/MySQL/thrift client, but it wasn't clear that doing so would have much performance win, as jruby itself wasn't dramatically faster than MRI. It would have, however, made it really easy for us to offload intense bits of code to java code, which probably would have been a faster upgrade path than rewriting in scala as we did.
JRuby was easy because you can Maven require it from a Java project. Ruby already has a Thrift IDL parser, so I just stole the AST from that, and used an ERb template to write out corresponding scala. The whole thing was maybe 200LOC.
But yeah, that's the only JRuby that ever did anything production related at Twitter.
We cut down our EC2 instance usage by 2/3 with more improvements yet to come. One machine alone can handle 1000 API calls / second - and our API calls are performing complex calculations, not just disk I/O.
It also allows us to deploy our API within customer's networks if they choose, which we previously accomplished using Virtual Machines - which sucked.
The article mentioned that due to the asynchronous nature of goroutines, instrumentation and metrics publishing was not a problem. I do not think this just applies to Go, but any primarily async languages -- like node.js, etc. However, this is still not close the JVM's incredibly sophisticated support for metrics and insights for your running applications.
Re: Go, I can see the value in having concurrency built-in, but I see http://www.paralleluniverse.co/'s libraries like Quasar and Comstat ultimately becoming the defacto standard in modern day Java programming.
With Java8 in rapid adoption and the upcoming changes in Java9, alongside pragmatic languages like http://kotlinlang.org with incredible tooling out of the box, Java is looking like a mighty fine eco-system to get started with.
That being said, it is extremely frustrating developing on a mature and advanced eco-system. As an engineer transitioning to Java from Python, I constantly have a lingering feeling that I'm not going to deliver idiomatic code and I don't know the best way to do things. Since I'm not aware of why design decisions were made in a certain way, probably to preserve backwards compatibility, I always ask myself why this was the best way and it's a bit difficult to research why. I also needed to familiarize myself with various patterns and Java's idiosyncrasies (1 public class per file, for example).
My 2cents.
Go is a pleasure to work in. Java, less.
Sometimes I wish I'd stopped at Go instead of then exploring Racket, Ocaml, Clojure, Haskell, and Scala so I could continue to hold your opinion.
The end result still matters the most to me, but being so aware of how languages are getting in my way and making me do busywork is exhausting.
Go is used in fewer places, but those places seem to be very interesting/exciting/dynamic places to work for. My kind of places.
YMMV and all that...
It's also the only one that's not maintained by MongoDB Inc. Coincidental? :)
PS: And yes, `mgo` by Gustavo Niemeyer is pretty incredible.
import Database.MongoDB
import Control.Monad.Trans (liftIO)
main = do
pipe <- connect (host "127.0.0.1")
e <- access pipe master "baseball" run
close pipe
print e
run = do
clearTeams
insertTeams
allTeams >>= printDocs "All Teams"
nationalLeagueTeams >>= printDocs "National League Teams"
newYorkTeams >>= printDocs "New York Teams"
clearTeams = delete (select [] "team")
insertTeams = insertMany "team" [
["name" =: "Yankees", "home" =: ["city" =: "New York", "state" =: "NY"], "league" =: "American"],
["name" =: "Mets", "home" =: ["city" =: "New York", "state" =: "NY"], "league" =: "National"],
["name" =: "Phillies", "home" =: ["city" =: "Philadelphia", "state" =: "PA"], "league" =: "National"],
["name" =: "Red Sox", "home" =: ["city" =: "Boston", "state" =: "MA"], "league" =: "American"] ]
allTeams = rest =<< find (select [] "team") {sort = ["home.city" =: 1]}
nationalLeagueTeams = rest =<< find (select ["league" =: "National"] "team")
newYorkTeams = rest =<< find (select ["home.state" =: "NY"] "team") {project = ["name" =: 1, "league" =: 1]}
printDocs title docs = liftIO $ putStrLn title >> mapM_ (print . exclude ["_id"]) docs- the "one-process-per-request" meme along the post applies only to some ruby app servers (there are event loop and threaded models too, think thin, puma, passenger in some modes) and I guess reading between the lines that it's mostly a problem of thread-safety and async support, because of the gems Parse used to have, right? I'm sure that limits options at some point anyway, but the statement is misleading and not really explained, I'd love to hear more details
- I don't understand how the comments in the little Go file snippet applies in any way to "ruby" ; it may be rails caching mechanisms, or a specific gem, but I have a hard time mapping those very specific details to something intristic to ruby, it seems more like grumpy ruby bashing, like you'd have done php bashing 5 years ago
As all rewrite stories, I think there's a part of envy/excitement over the new cool tech you want to use (and that's fair! pleasure give you huge productivity boosts), and also a part of success related to the fact you know the kind of things you failed in the first version, so you won't make the same mistakes the 2nd time.
I'd love to hear finer details on those points! Great article overall anyway
I'm hoping to get some followup posts from the backend eng team on specific interesting problems we ran into during the rewrite.
& yes deploys with go are the freaking bomb :)
We serve millions of requests per day and have some slow responses (75+ ms) but on any given day our servers handle 175 requests per second without breaking a sweat. =/
So for example if an app does something bad like performing a full table scan on every request against a 300 million collection, so every request to that backend is timing out at 30 seconds, and there are thousands of them per second, well -- pretty soon your fixed pool will be full of requests timing out to that backend.
Many organizations scale out too soon or for the wrong reasons, and some just have ineffecient database queries and other things that result in bad request times that they could also optimize -- which helps the user regardless of scale with faster page loads.
But of all problems to have, there are many worse than exploding growth.
Of course, on the other side of things, everything feels rosy -- but counterfactually all the effort they spent on this could have potentially gone elsewhere if they resolved their scalability issues with Rails in a simpler matter. (Or even better, contributed those solutions back to the community.) This was a move that fortunately worked out, but it sounds pretty high risk to me and is the type of thing that can kill companies if they bet wrongly.
Could you elaborate on this? It sounds a bit scary. Does this mean that Rails tries to decode a URL several times until it can't be decoded? If so, isn't this problematic if some (arguably crazy) person tries to send "%2F" literally, not "/"? I'm half sure I'm misinterpreting, so here to ask.
Like Go using say warp[0] and async[1] likely with similar performance numbers but with less code, more static typing guarantees, and simpler[2] code. Like Go though, you'd deploy with a static binary. This is just a wild guess though, I would need to know specifics of the Go application they've created.
0: http://hackage.haskell.org/package/warp 1: http://hackage.haskell.org/package/async 2: I find redundant code like Go requires[3] to be more complex. Haskell code can be complex, but simple, straight-forward, not trying to be complex Haskell is very simple and concise. 3: Well, you can use interfaces and lose type safety. Or you can use reflection and make things dog slow.