Twitter Shifting More Code to JVM, Citing Performance and Encapsulation
infoq.com
infoq.com
Neither way forward was presented as an obvious win. Building something that grows meant overengineering and trying to forecast the future. Building something that evolves meant constantly rewriting things, taking time away from new development.
Twitter seems to have taken the "evolution" route, and it has served them well. I'm not at all sure they would be this successful if they'd tried to build on the JVM right from the start using the technologies available at the time.
To this day it takes real effort to convince Java programmers that a lot of "best practices" are anything but.
For instance I still see Java programmers use frameworks that require them to maintain mountains of flimsy XML configs for things that they ought to have done in code. Both to get the benefit of letting the compiler do the work of weeding out boneheaded errors and to get rid of unnecessary flexibility that just leads to more work and more confusing, hard to read code.
Java is a great language in which you can be very productive. But productivity means that you have to crack down on people who drag along J2EE-crap, or whatever crap was invented to make J2EE-crap slightly less crap. It also means you have to mentor people actively and harass anyone who even thinks of doing in XML what can be accomplished much faster in testable, compiled Java code.
Indeed. Back in 1998 I worked on a Java Web framework that was quite slick, easy to use, and relatively lightweight.
It used URLs to map requests to classes and methods (e.g. /foo/bar/baz meant get a reference to an instance of class Foo and invoke bar("baz") ). It used convention over configuration for the common stuff (e.g. where to find classes and views) and was quite robust using a custom request proxy to handle the interaction with Apache.
It was easy to add new behavior (while the app was running, even), easy to do internationalization and custom styling, and easy to deploy.
When the company I was working for got bought the owners decided that they wanted some J2EE thing that could be maintained by a redundant array of mediocre Java hacks so they dumped the fast, efficient, and fun custom code and rebuild the app using beans and entities and war files and all that nonsense.
I quit shortly thereafter.
Since then I've seen countless people bemoan the state of Java Web development, but the truth is that much of the sorrow is self-induced. You don't have to build complex, complicated Java frameworks. And you don't have use only what others have put before you.
That is a problem with Java, but it is not the problem with Java.
The JVM platform, running server-side, is a proven technology that lets you get stuff done. The Java language, however, is underpowered by today's standards. It has essentially no technical advantage over newer JVM-based alternatives, which can do things better by taking advantage of a decade or more of hindsight that Java never had. Meanwhile, it does have clear technical limitations that make it difficult to use some important programming concepts.
For established projects, Java-the-language survives based on inertia, of course. However, I suspect that for newer projects, some of the alternative JVM languages are now mature enough, both technically and in other important areas such as community and tools development, that using them instead of Java itself should be almost automatic for a lot of projects.
The Java language does lead to unnecessarily verbose code, but not shockingly so. I'd prefer somewhat verbose but idiomatic Java code over similarly idiomatic code-bases in many other languages.
Of course, part of the problem is that there are far too many bad programmers (and designers) in the Java community and the reputation of the language suffers because of these people. Java even started out with some of these horribly bad programmes on the staff of Sun. (Those of you who actually read the source code of the standard libraries can probably guess who I am talking about).
The thing is, though: I don't really see this as a problem. In part because it is possible to write reasonable code in Java itself and in part because there are alternative languages that are interoperable with plain Java such as Scala and Groovy. This means you can benefit from the Java ecosystem, but you have more choice when it comes to how you write code.
(That being said: I am skeptical to introducing some things into Java because I don't think it will lead to better code. For instance I don't want closures because it will tempt inept programmers to over-use these features in contexts where they are not a good fit. Closures can turn relatively linearly readable code into a horribly tangled mess that can be hard to reason about. And as I said earlier: Java has attracted a lot of talent-free programmers)
I guess that's where we'd have to agree to disagree. It sounds like you are concerned that providing more expressive features, such as closures, will tempt bad programmers, of which the Java community has plenty, to write bad code. I take the view that there are also plenty of competent and some very good programmers out there, and that Java does not let these people get much useful work done as easily as more modern alternatives.
For the competent people, Java is horrendously verbose, and the evidence suggests that this does have a significant impact on the pace of development and the robustness of the finished product. Google Scholar will find you a wealth of reports about this if you're not familiar with the research, mostly based on controlled and fairly small-scale academic experiments, but in some cases looking at industrial case studies as well.
As you suggested, the solution to this is often to use other languages that also run on the JVM in preference to Java itself. I just don't see what redeeming value Java has at that point, unless your dev teams are composed of lots of incompetent devs, in which case frankly you're in trouble whatever language you choose.
This isn't (only) because I dislike typing a lot but because I try really hard to a) par down the problem until I understand it in its most basic form, and b) try to express any solution as simply as possible and as readably as possible.
if mere code terseness is important to you there are other languages that will provide that. however, there is no language that will compensate for the inability to avoid implementing lots of unneeded stuff that comes from not understanding the problem. and for the most part: that's what separates the really good programmers from the bad ones and this is where the majority of the line-count gets spent.
that being said, I fully agree that Java is too verbose. there are a lot of things that it'd be nice if Java took care of for you. this is what I like about, for instance, Groovy. (I've only dabbled in Scala so it would be premature to sing its praise, but I am sufficiently impressed to have planned to do a project in Scala this fall).
I'm not advocating a very terse programming style, as tyically seen in Perl or C code.
However, Java doesn't let you compose basic data processing easily. If you have a list of Xs and you want to find those fitting a certain criterion, just compare the effort you have to go to in Haskell:
results = filter (<10) xs
and in C#: var results = from x in xs where x < 10 select x
and in Python: results = [x for x in xs if x < 10]
and in Java with someone's collections library on top: results = CollectionLib.filter(xs, new CollectionLib.BooleanTest<X>()
{
public boolean test(X x)
{
return x < 10;
}
} );
That's pretty typical of the problems with pure computations, in my experience.When you start talking about I/O and other programming with a time dimension, you find similarly clunky handling of everyday things like events/publish-subscribe/observer/whatever you want to call it, not least because everything has to be in an explicit interface before you even start. Meanwhile, other languages these days are offering tools like message passing and actor models, which both scale up with system size without compromising the basic architecture and support concurrency with relative readability and safety.
Please notice that none of this has anything to do with terseness as such. It's about whether the language provides simple, readable tools to implement widely applicable programming techniques.
kill me. Makes me want to kill myself.
All the XML-wrapper tools for Java are supposedly the reason why there is so much XML config stuff going on - but the result is that you end up writing soo much code in XML instead of the actual language that you're supposed to be using.
It's like somebody looked quickly at the MVC pattern, got it all wrong, and now Java has very tight coupling between XML and Java code for any production environment that hopes to leverage existing tools.
Jave EE, on the other hand, is crap.
I have worked with one of the authors on a "serious project" in a bank and over time we gradually ripped out the frameworks, making the code much more explicit and less magic.
I guess it depends on the team you are working with. If you have solid, experienced engineers who understand the business domain, using Spring may not be relevant. But if you have a team that have different levels of skills and experience, a framework that abstracts away some of the underlying complexity is useful.
If you are building a real time trading system, you probably don't want to use Spring because you want to control each and every component. But for most other projects it can be useful.
Btw - your link mentions TDD. What tools did you use? (just out of curiousity..)
I worked with Java EE in 2004, took a break, and started working with it again last year. It's not too bad, except for JSF. JSF is really a horrible mess. They should scrap it and start over.
Java may be a bit limp. But there's no reason to hit it over the head with a bent crutch as heavy as Spring. That doesn't make it better.
Generally.
All of the XML config stuff was, to some extent, what turned me away from Java so many years ago.
If I were to start over today, what kind of frameworks would be the wiser choices for desktop apps (gui frameworks) or web apps (web frameworks, maybe an ORM)?
But if I were to come up with an answer it would be something along these lines: don't use of a framework or tool that cannot, with relative ease, be replaced by something else. For instance, in a well-designed networked application it is usually simple to replace one networking library with another. Or to replace one HTTP implementation with another. And I am not really talking about drop-in replacements.
Years ago I spent about about a week replacing the entire networking layer in a high traffic server, going from a pre-NIO blocking design to an asynchronous NIO-based design. This actually changed the entire execution model of the system as well as the networking parts, but it had a lot less impact than you'd think because there was proper separation of concerns, strong non-leaky abstractions and a lot of code that was written by people who were disciplined and consistent designers.
As for GUIs, I am not really the right person to ask. I've dabbled a bit in GWT, but I am not entirely certain I like it. Perhaps not so much because of GWT itself, but because it is tiresome to deal with building and deploying.
ORMs are generally a bad idea. Avoid them. You will feel some initial thrill when you can do some simple magic tricks and then everything ends in tears when you find that you actually have to understand exactly how it works and dig into the innards. Definitely not worth the trouble. (If you use ORMs by way of annotations you are doubly fucked because you will have one more thing that can go wrong which then necessitates dipping your toes into territory that you are not dealing with on a daily basis. I have no idea where some developers find the guts to depend on complex yet fragile subsystems that they have zero understanding of)
Instead you should design internal application specific APIs for dealing with stored state.
For instance, if you are writing a blogging server, you should design a interface that provides the operations you need against the blog store. Start by just implementing the storage operations in an implementation class. Then, if you need support for different types of blog stores, you extract an interface definition and then write implementations of that interface. (Of course, when you write the first implementation class you keep in mind that you might want to turn it into an interface later. This should keep you honest and ensure that you never, ever leak types that are specific to the underlying storage through your API).
In one of my current projects I did just that: I created an abstraction over what I needed to store in a database. The initial prototype didn't even use a database -- it was backed by in-memory data structures. Lists and Maps. This allowed me to prototype, experiment and discover what I actually needed without being side-tracked by details on how to realize this in a database.
Eventually we wrote implementations for both an SQL database (mostly as an experiment) and Cassandra. At that time, people depending on this server had already integrated with it -- before it was even capable of persisting a single byte to disk. As I wrote the in-memory implementation I wrote extensive unit tests. Both to test for the correctness of the code, but also to document what behavior was expected of an implementation. Not only did we later apply the same battery of unit tests to the other implementations, but the unit tests became the measure of whether new backends would be compliant.
As I said earlier, it is hard to give general advice, but I think it is very important to learn how to design software rather than picking a framework that will dictate the design for you. It is very hard to undo choice of architecture so at the very least one should make an effort to learn how to think about, and design, architecture. If nothing else so you can later choose the Least Evil Alternative.
I think 90% of people who got on the J2EE bandwagon were clueless about architecture and just did as they were told. the remaining 10% may have cared about architecture, but were not sufficiently averse to complexity and mindful about programming ergonomics to realize what a horribly bad idea it was. Of course, by the time people realized J2EE was a waste of time they had all this value locked into code that was really, really hard to re-use in a different context.
As for Spring and the over-use of dependency-injection and autowiring, that too will pass once the loudest monkeys in the tree get to change jobs a couple of times and realize that breeding complexity by scattering knowledge across a bunch of files is not a terribly bright thing to do. People usually get to hate Spring once they inherit someone else's non-trivial Spring-infested codebase.
I happen to not like C++. So I don't use it. See? Easy.
I'm not sure why you think that, or why you think Twiiter had a 'grow' vs 'evolve' strategy. Their approach seemed not nearly so self-aware, and resulted in serious, core deficiencies that required considerable time and expense to correct. From an external position, it very much did not look like a rational considered approach to solving he problem through 'evolution'.
All evolution strategies involve serious, core deficiencies, that's the trade-off, that's why it isn't 100% obviously better to evolve than to grow. But I do suggest that evolution--however it came about at Twitter HQ--has worked out for them.
It allows you to focus on features and pivoting over ceremony (boilerplatey / architecture stuff) and when you finally start hitting those performance limits, you have a luxury problem.
It will also be more apparent where to invest time spent scaling your product, whereas if you try to optimize out of the gate, you may still hit unforeseen performance issues.
This always sounds nice, but I've found that ceremony almost never really reduces productivity by much. The real wins/losses almost never have to do with ceremony related features of language, but rather architectural and framework components.
You build something that evolves by bringing in the smallest architecture and least fx components and then build. The language you use may dictate the fx components to some extent -- otherwise it's typically just to make developers feel happy (which is important, but really is just about morale more than anything truly inherent in the productivity of the language).
I was arguing in favor of 'growing something' as opposed to 'designing something' with regards to a hopefully growing userbase.
Language choice may be of lesser importance as you say, but I do feel that some languages fit the growing strategy better while others have a more design up front feel to them.
on the other side the reality of a startup is you're going to push out so much code / features / product based on demand so fast that a lot of the things that you consider building for the future will be thrown out the window pretty fast. Just build it, see if it works, and move to the next thing.
It confused me, because Brooks talks about "growing' software (as opposed to building/planning it), which is I think what you mean by "something that evolves".
FWIW, I think the standard wisdom today is to evolve (your term) software (e.g. agile ideas of YAGNI, DTSTTCPW), and only to scale/plan it if you have a very clear idea of what you're doing (e.g. frozen specs, which exist in some government/military domains; long-term standards; mathematics) - and you also know how to do it, having done it a few times before.
The key is, when designing your growth architecture (and I don't consider it overengineering--a few simple decisions allow you to scale pretty well to a fairly large load), is to code to interface rather than code to implementation. If things evolve/change over time, you've already got half the work done for you because things don't spontaneously break on change.
However, at the time it might have gone the other way: Perhaps their user base and volume might have increased at a slower place, leaving plenty of time to evolve their scalability, while their user experience might have required relentless change.
In which case, they should have planned for their user experience to "scale" rather than their infrastructure.
A scalable solution is usually a more rigid solution as well, and that can be a problem for a startup that may need to try a few ideas before they hit on what works.
It takes async network programming with netty into a functional programming paradigm. Programming scala/finagle network services is much nicer IMO than coffeescript, ruby/em/fibers, raw netty. I can't wait until we release our finagle-based cassandra client. It's been really nice to work with.
Here's some sample code:
https://github.com/twitter/finagle/tree/master/finagle-examp...
Also, I may be in need of a nice async cassandra client very soon, any chance of an alpha preview? (e-mail in profile)
"To allow developers to choose the best language for the job, Twitter has invested a lot of effort in writing internal frameworks which encapsulate common concerns."
The single best investment companies can make is to allow developers to choose their specialty, and their language. Otherwise you have a huge overhead of skill set mismatch. And your talent pool can be bigger if you're open to more than one language.
It's a refreshing view.
In my personal experience, for reasonably isolated and small components, 80% of time spent creating a component of code is spent on making decisions. By the end of it, when all the decisions have been made, I could probably rewrite the entire thing from scratch in 20% of the time—and it would likely be better in every way.
It would seem that the first iteration is largely prototyping. If you can save a significant multiplier of time by choosing a different language than the rest of the stack for the prototype and potentially rewrite it later if necessary, why not? Perhaps by the end of it all, you'd break even on time but end up with multiple implementations and better code.
It seems unlikely that 10 developers in a startup wouldn't be able to maintain code written in a few different languages.
The way you keep a codebase easy to manage is to divide it up into small projects that you can "finish". When was the last time you hacked on glibc? Never? That's how parts of your infrastructure should work: get them right, then forget about them. Using the best language for the job makes this significantly easier.
This, imo, is part of having developers choose the best language for the job. They chose ruby to get their mvp out quickly and be able to grow their userbase. After that, they pinpointed trouble areas in their architecture and made pragmatic choices in fixing up those areas.
While I suppose it's possible to "let developers pick the best tool for the job" and end up with 5 different languages in your stack, most good developers are likely to steer clear of such an endgame, unless the trade-offs are very, very clear.
If you can allow multiple languages to share common code like you can on the JVM, then I say it's ok to go crazy.
The primary driver is honestly encapsulation, so we can iterate faster as a company. Having a single, monolithic application codebase is not amenable to quick movement on a per-team basis. So when we decide to encapsulate something, then because of our performance concerns, its better to rewrite it in the JVM for most systems, than to write a new Ruby system.
It sounds from that like their primary driver for using the JVM is actually performance, but that they only decide to rewrite components when encapsulation drives them to do so. I can't see how the JVM provides any encapsulation benefits over Ruby for new systems.
A 'productivity boon'? I don't understand. At the risk of invoking the ancient static vs. dynamic religious war, this statement makes no sense to me.
I get that if your codebase is tangled enough, and your unit test suite is inadequate to "guarantee that your dataflow is more or less going to work" that maybe rewriting significant portions of it in a type-safe system makes sense. I guess. But without specific code examples it's hard to say exactly what he's talking about.
Myself, I've spent many years in both static and dynamic environments and I know exactly where I'm more productive -- and it's not wrestling complex parameterized types to the ground, pulling up abstract classes or interfaces, and/or configuring IOC containers, abstract factories and the like.
I wonder though -- this has echoes of Alex Payne's criticisms a couple of years ago, which I think Obie Fernandez addressed pretty well:
http://blog.obiefernandez.com/content/2009/04/my-reasoned-re...
I don't know what they mean either, but my first guess has to do with company size and mobility of staff.
I love dynamic languages most when I'm coding solo or with small teams. I don't need to express a lot of things to the computer when they're so clear in my head. But if I'm going to take over an adequately maintained code base, I'd rather it be in a static language, because more of the intent is explicit.
At this point Twitter has a lot of engineers and is still growing, and they're in a very dynamic business. It's plausible to me that they get a global productivity boost even though static languages could feel like a productivity hit to each individual engineer.
Out of interest would you feel the same if both codebases had adequate test coverage?
In the post one of the reasons given was that with a static language you can pretty much guarantee that a dataflow is going to work, that you won't be caught out by getting the wrong type. I'd see this being most useful at the edges of the system, and in those cases incoming data would normally go through some validation anyway (including through a schema in many cases) which would normally make clear the types involved.
Having said that I do think in those cases being able to specify types can makes things slightly easier for newcomers, I'm just surprised its seen as a big enough advantage to be one of the key motivators to switching language.
No project I've worked on in twenty years had test coverage I could call "adequate", though I realize this is partly my fault. Hard-core TDD from day one might get you as far as "mediocre", and the industry average is much worse than that.
Your unit tests can never guarantee anything about the dataflow of your system. That's not the purpose of unit tests. And unless you have a meta test system, that tests the properties of your unit tests, you can't get system-wide dataflow information.
static typing might be productivity gain when more persons are involved and dynamic typing might be productivity gain when less persons are involved
Or it might not matter.
A big country needs very big laws, but they are never strict (static), nor dynamic.
A small country, a town might not need laws at all (they are known - dynamic).
Not a very good analogy, but still...
Pragmatic failure inevitably leads to analysis paralysis. Just worry about getting stuff done. :)
Neither may be related but for a large company with very little product they seem to produce astoundingly little.
Do they still have these problems or in these aspects python is better than ruby?
The simple Web Server -> ORM -> Relational database architecture that most modern web frameworks utilize can easily break down under tremendous concurrent load, especially if you attempt to run it on commodity hardware.
IMO, the problem is that Twitter was Ruby's poster while both were still in ascension (might as well say they still are).
The first migration they did was porting their message queue from Ruby to Scala. As mentioned on a TheGuardian article, they migrated because Starling was crashing too often and dropping tweets, so they had to stop the site and migrate manually. Also, another member claimed performance problems on some blog posts. However, Starling doesn't perform that well to Ruby standards either, so it's very difficult to say.
I'd say the migration was for "social" reasons. The team is clearly waaaay more comfortable and enthusiastic about Scala, also they seem to prefer Static typing, so using it is the best choice indeed...