Lessons learned porting 50k loc from Java to Go
blog.kowalczyk.info
blog.kowalczyk.info
I’ve worked with C# for a decade, and I’ve yet to see a use of inheritance that wouldn’t have been easier to maintain in the long run without using inheritance. We’ve limited our own useage to override methods in the standard library, but even then it’s often used to implement things that are really just terrible practices. Like adding search functions for AD extension fields or increasing the timeout in one of the older web clients.
Many languages (thinking c#, Java, c++) say they promote composition over inheritance but yet make inheritance so mmuch easier to use.
Since the 90's, any good CS book about OOP paradigms had discussions about is-a and has-a and how to make the best use of each, depending on the desired application architecture.
The thing is, such books usually aren't taught on bootcamps.
When we have interns, they’ll sometimes build things with inheritance. So it’s certainly still a thing.
I’ve yet to see a real world use of it, where you wouldn’t have been better of not using it though. My real world is the relatively boring world of enterprise public sector software, however, and maybe I’m simply oblivious to where inheritance might be worthwhile.
Initially, VBX only allowed for composition as well, COM introduced interface inheritance with delegation, when one wanted to override a couple of methods, but not the remaining several ones.
And now UWP offers mechanisms to do implementation inheritance in COM, because everyone got tired to write delegating code for is-a relations.
Inheritance and composition are both tools, it is up to each one to learn how to use them appropriately.
Am I missing something?
Looking at the history of inheritence a few languages I know:
- C++: multiple inheritance super confusing expressions and understanding of things like diamond inheritance. Initialization becoming a really complex thing to understand.
- Java: learning from the mistake of C++, determines multiple inheritance is bad, only allows inheriting data from one super class and then API (and now in 1.8 functionality) from any number of interfaces.
- Rust: realizes inheritance of any data is bad, and only allows for any number of Trait's (interfaces) to be implemented on types.
There's a difference, and ultimately it's that you're not carrying data via the inheritance chain, only functions. The benefit here being that it's much easier to reason about a type and what data it has that determines how it works in the broader system. This allows for very explicit inclusion of data from other types, as opposed to implicit inclusion with C++ and/or Java. In the end moving away from data inheritance leads to fewer mistakes and easier understanding of the code.
A: Java coders.
Inheritance defines a mechanism that is sometimes useful. It is useless as an interface design tool.
I didn’t mean to suggest that it’s idiomatic in C++ to create strange inheritance graphs, anyone who’s done it once realizes it’s a bad pattern. The issue is more that the language -allows- it.
IMHO - There are many managerial short comings in software development that lead to abuse of legacy design. OOP by design allows you to abstract and ignore not only original implementation decisions but reflections of those decisions in the architecture. Polymorphism makes sense when it's the classic ICat: IAnimal example but so often it _becomes_ IHouseFly : IAnimal because all the contracts expect IAnimal and the deadline is.....tomorrow.
I personally don't have a good solution given the pragmatic counter argument that the ROI on a system which is cheap to develop but is patched over 5 years may be equal to or cheaper than one that is expensive to develop and _still_ needs to be patched over it's 5 year lifetime. Let's call this the used car problem. A new car is more reliable but now that used cars are reliable _enough_ it's harder to convince people _not_ to gamble with a lower upfront cost.
I absolutely _love_ Rust and the code I'm written feels bullet proof. No idea how long it would take a team of my C# peers to be even 1/10th as productive in Rust/Go as they are in Visual Studio and C#.
Something that means rework, sometimes you end up with a parent method which is overridden in every child, possibly because each override was spread out over a long time period and no one bothered to look outside the child.
Inheritance makes sense with interfaces in C# but for the rest I think composition is just a better way of sharing.
There are other languages that have better support for composition, while still maintaining support for inheritance (e.g. Kotlin). Inheritance has its uses, as evident by the fact that Rust is considering adding support for inheritance (e.g. writing GUIs).
https://play.rust-lang.org/?version=stable&mode=debug&editio...
A lot of stuff in Java is just because (by now) it has a lot of history and baggage. JavaBeans were kind of cool back in the late nineties but mostly it's a really annoying convention by now. The notion of accessors is not obsolete but unfortunately was not part of the language design when they created Java. Since Java has reflection, they came up with naming conventions that allow code to inspect object instances and do things with properties based on naming conventions. This has been key to a lot of the Java enterprise stuff that happened after this. And it also facilitated creating UI builders that came with IDEs like jbuilder, netbeans, visual cafe, etc. when having a UI builder was still a base requirement for an IDE. Eclipse sort of broke that tradition by not including one (initially).
These days UI builders are basically very uncommon; which makes javascript frontend work particularly tedious and repetitive. Because there are no conventions, no typing, and quite often no meaningful test coverage in JS, it is basically impossible to deal with that in a UI tool. So, accessors matter and are not obsolete at all.
Most modern languages have more elegant ways to do essentially the same type of accessor mechanisms in the language. And also languages like Smalltalk had this (as well as UI building tools, refactoring, and a lot of other stuff that is still science fiction in the javascript world).
If you look at Kotlin, they fixed this while retaining the ability to expose code back to Java.
For example kotlin has properties with accessors that you can optionally override. Normally you just type val foo="bar" and you have a string property with an inferred type of String. The setters/getters generated under the hood and used automatically when you assign or use the variable. If you want you can customise the accessors or use something called delegated properties that e.g. turn a property getter into a function or use lazy intialization. Once compiled, a java class that uses that code would see the normal setters and getters as if it was a normal Java class. Likewise when accessing Java code from Kotlin you use java properties as if they were normal kotlin properties (i.e. without using setFoo(foo)/getFoo()). This makes Kotlin a really nice way of using legacy Java code.
The cool thing is that you can convert to Kotlin on a file by file basis.
After that you get to the idiomatic stuff like e.g. making properties read only getting rid of multiple constructors by introducing default parameters. Getting rid of the builder pattern (mostly redundant in kotlin), introducing data classes where that makes sense, using lateinit vars to make nullable vals nonnullable, etc. Technically you are at that point improving things.
A new coder diving in to the code will need to either know all the Lombok annotations by heart or look up what every single one of them does.
Also it makes debugging kinda crappy for the same reason.
you can also delombok if you ever get tired of it, and now your back to manual.
I personally don't use lombok, as i'm not offended by the verbosity, given that ides since forever have done all the work, but if it's something you are bothered by, well. that.
If an argument is that javabeans are the wrong pattern to use, that's a different argument, and unrelated to lombok.
But he never talks about motivation of the original code. Like why it was generic in Java. Sometimes generic code in other languages can be ported to Go to two or three functions, because the general use of the generic code wasn't as generic as the designer may have envisioned, or it just became a nail to their hammer.
Like was the inheritance really important, or could it be implemented another way in Go?
But I guess if you're not questioning your client's motives, you may not be fully questioning how something ought to work, and instead just ensure it works as it does right now.
It's possible they just needed a Go-library for their client code, and then the client code can be tweaked later on to be more idiomatic Go.
That being said, from a business point of few, making Go code base be as close as possible to Java version makes a lot of sense.
Java client is written the way it is. It already shipped and is used by people. Breaking its API was out of the question.
RavenDB is 10 years old and will most likely be here for the next 10 years.
In those future years, both Java client and Go client will have to be evolved to provide access to new capabilities of the server.
The closer the two code bases are, the less effort it takes to maintain those code bases.
It requires database drivers (client libraries) for as many languages as possible.
More client libraries, more programmers can use RavenDB, more licenses for RavenDB sold.
They already have C#/Java/Python/Node.js libraries.
I ported Java client library to Go, so that people who program in Go can access RavenDB database.
As a result, the original and most featureful client is for C# / .NET.
Java client is a port of C# client, done in house by Hibernating Rhinos (the creators of RavenDB).
As far as I can tell, other clients (Python / Node.js and the Go client that I wrote) were contracted to outside people.
The company suggested starting from Java code base. It makes sense because C# client heavily uses LINQ, which is unique to C# (neither Java nor Go has LINQ-like capabilities).
I didn't dig much into non-Java clients so can't speak much to that.
Overall, I was surprised how similar I was able to make Go code to Java code.
Changing from exceptions to errors was pervasive but a simple, mechanical transformation.
Porting functions using generics was the biggest hurdle.
Porting functions that use overloading was easy but annoying.
That being said, a Java code base that heavily used virtual functions and deep inheritance hierarchy would be more challenging to port to Go. Lucky for me, this code wasn't.
Not true since Java 8, with the introduction of streams and functional interfaces, which keep being improved with each release.
What Java doesn't have are expression trees, which are convenient but not a requirement for LINQ like features.
The objective of the port was to keep it as close as possible to Java code base, as it needs to be kept in sync with Java changes in the future
so it doesn't look like they are ditching their Java client.
There needn't be any "problem with Java". This is a client library. You want to have it in as many languages as possible.
Performance is also usually on par - unless java's equivalent is built around a lot of reflection (orm and what not).
However, go compiled binaries are usually smaller, and code IMO is much more readable.
For example, Go's err != nil pattern is often cited as being ugly, but good go code will often remove errors by design.
There's a good post by Dave Cheney about this; https://dave.cheney.net/2019/01/27/eliminate-error-handling-....
I think this is equally true of Java. Most Java code I've seen disgusts me, but I've also seen some beautifully written pieces.
Another thing is that in many cases there is no good sentinel value to return that naturally leads to exit from loops or complex logic to check for error at the end of a function.
Go codebases do tend to be a little shorter due to lack of getters/setters... and generics are not used that often in production codebases anyway, relative to all the rest of typical code (that’s procedural anyway).
no.
They're usually longer. See: no generics, error handling, among other things.
If you want to argue 'usually' then you could consider all the build scripts and XML files, class boilerplate and exception code of Java to be 'usually longer'.
Please don't troll.
Code duplication, just one example: https://godoc.org/github.com/cznic/mathutil#Max
Even with exception handling code, Java is shorter. Build scripts are not part of this discussion.
Another reason for the difference is that the Java client is actually considered to be the primary one, which all the other ones are based off.
The Python client is missing a fair number of tests, though.
Expect a python program to be half or a third of the size of java/c#/go.
Regarding the author's frustration of moving POJOs over (and their variable declarations), I used sublime text to select all variables based on a shard token, then cut them and moved them over by word. You can then lower case the first letters of every word, and then find-and-replace by type using a shared token. I found this method very quick and effective.
>
>Go does not.
Huh? As a package maintainer for a lot of Go stuff, I had to deal with tricky cyclic dependencies several times, especially in Google own Go packages like golang.org/x/build and google cloud.
More seriously, they probably means find ways to restructure packages and extract shared dependencies into subpackages to avoid cyclical dependencies.
Packaging for a distro requires building Go packages isolated from the network in a chroot and having all the dependencies previously built in the same manner. So it is an iterative process, building blocks by building blocks. If you have a cyclic dependencies, you can't build iteratively, and you'll have to excise certain part of the code to eliminate the cycle.
I would survey existing libraries for other JSON document databases (MongoDB, Google's Firestore etc.) and steal all their best ideas.
But honestly, good API design is something that needs a lot of time to reach the polished state.
Whatever I can come with today, I'm sure a year from now I would find ways to improve that.
That has been my experience with much smaller libraries I wrote.
The more time you get to spend thinking about the problem, the better solution you can come up with.
It's just it's nice to be able to track code coverage over time, which is what codecov.io provides.
Also, running all tests (to get full code coverage) takes 20 minutes and makes fan on my laptop unhappy.
The way it works is that on every checkin the CI job runs all the tests with code coverage enabled and uploads the results to codecov.
Codecov can then plot coverage over time.
My gripe is with inaccurate accounting of empty lines (like comments or struct definitions) by codecov. Go's tool to visualize this count them properly. I don't know if it's codecov or maybe I'm not sending the data properly.
> Codecov is barely adequate. For Go, they count non-code lines (comments etc.) as not executed. It's impossible to get 100% code coverage as reported by the tool.
I have an open source Go project that has some comments in methods, but still achieves 100% coverage using Codecov.io — I'm not sure what I do differently to yourself? (Perhaps I'm not using any inline struct definitions?)
Here's a link, in case there's anything useful to you in my .travis.yml ? https://github.com/jimsmart/store4
HTH
I don't know, maybe I'm not counting things right but for example https://codecov.io/gh/ravendb/ravendb-go-client/src/master/d... shows less then 80% coverage and there are only 2 lines not exected out of at least 18, which should be at most 10% counted as not covered.
Codecov* doesn't count an 'if' statement as having full coverage unless one tests both outcomes: so the yellow lines here have been executed, but do not count towards your coverage score.
Granted, one could argue that that's not very generous! But on the other hand: those yellow lines have not been fully tested, despite being executed, so I can understand their decision.
In the linked code, just implement a couple of simple tests to test for the expected error conditions: it's easy (here at least) and ensures the code behaves as expected. (Obviously not all partial/no coverage lines will be so easy to hit with tests, it might not always be possible to easily get 100% coverage, but hey: start with the low hanging fruit!)
* I say Codecov here, but I highly suspect that they may simply be using Go's coverage reports under the hood?
I would rather go with CirrusCI which has windows, macos, linux and freebsd support, is much faster and is easier to work with. appveyor being the slowest, and travis having the worst features.
func (q Query) GroupBy(field string) Query {...}
PROCEDURE (q : Query) GroupBy*(field: String): Query;
BEGIN
(* .... *)
END GroupBy;Note that there are a few odd bits to Go methods: the receiver can be a value or a pointer, and if the receiver is a pointer it might be null, because method calls on concrete types are statically dispatched so
type S struct{}
func (self *S) Foo() {
fmt.Printf("%v\n", self)
}
func main() {
var s *S
s.Foo()
}
will print "<nil>".But RavenDB is 10 years old. Java client already exists and has 50 thousands lines of code.
If I was trying to rewrite that from scratch, I'm afraid Joel Spolsky would find me and spank me.
I'm not sure a tool could do 100% translation but I'm sure it could do a lot.
A surprisingly large amount of time was just moving the order of variable declaration from Java's "type name" to Go's "name type" and renaming, say, "String" to "string".
If a tool did that for me, it would save a ton of time.
Unfortunately, the upfront time investment to learn enough to write even the simplest translator would probably be greater than time saved on one project.
But depending on how the Java client uses null, "" can do just as well. It's not like you have many other options (except to add your own composite struct on top of String, or to use a guard value that's still a string).
The goal of the project was to enable those Go programmers to be able to use RavenDB database in projects written in Go.
The company also maintains Python client library and Node.js client library.
It's not about Go vs. Java as a technology but enabling as many programmers as possible to use RavenDB.
The bigger part might be transformation of the programming languages types into something useful for the client (e.g. through serialization), which has be be redone anyway for each language. And after that the question comes up whether sharing the remaining things yields enough benefits to justify the hassle of having a dependency which is less portable, requires another build system, etc.
That's my general experience with those kinds of projects - I don't know enough about RavenDB in particular to tell if it's the same here.
That said I'm still waiting for when the language ecosystems start to standardize how they interact with each other. That is a very hard nut to solve, but it would also be a quantum leap forward.
A feature that exists since around the mid-2000's.
It was only missing from the gratis version of Java, the large majority of commercial JVMs always had AOT support.
The goal for the company was to enable an estimated 1 million of Go programmers to use RavenDB database.
This is Java AND Go (and Python and Node) scenario, not Java OR Go.
Static typing?
Simplified toolchain without make/maven files?
My big one after swearing off of Java was never needing an IDE to do even the simplest things - is Kotlin useable without an IDE?
Yes. The JVM is Java's biggest downside, so why would you want to just move to a different language on the same overly complex (to put it mildly) runtime.
I have a rudimentary understanding of the Linux and BSD kernels and how they can impact certain parts of my applications. But I have zero knowledge of Windows. But I don't need to, because the JVM engineers do, and things Just Work (TM).
I'm not smart enough to learn the intricacies of every platform, and thankfully I don't need to be.
I guess you have similar issues with POSIX and the abstractions offered by high level programming languages? I assume you write some CPU specific assembly and understand all the possible failure modes of that hardware?
Yes, you do, it's called abstraction.
For example, C compiler adds complexity when compared with assembler, but that lets you significantly simplify your code.
And without abstraction, modern software development would not be possible.
Of course too much abstraction can be a problem (which is why some people still program in assembly), but that does not mean JVM will always be the wrong solution.
Kotlin native is coming along, uber jars are popular so only other departments is the JRE, Docker can bundle that
> Static typing?
Yes with good type inference and "null safety"
> Simplified toolchain without make/maven files?
You can use the compiler directly? But for anything serious a build tool is pretty essential for any language? At least you have options
> My big one after swearing off of Java was never needing an IDE to do even the simplest things - is Kotlin useable without an IDE?
You could write it with a Morse Code Keyer if you really wanted
Not only not an advantage, but kotlin does have that so I don't understand what you were going for here.
Kotlin can just use the Java client directly.