929 karma · joined May 12, 2012
TL;DR the ISO standards committees behind SQL are working on bring graph query languages and SQL databases closer together.
Obviously there is work to be done at the storage/query planning layer after that, but I’m hopeful once the surface exists more widely that will drive more work in those areas.
I was expecting some novel analysis of in depth data scraped from GitHub APIs, or the results of a qualitative survey of big Java shops, or … something?
Instead what we got was a reposting of a few surface level stats pulled from other blog posts which have tolled the death of Java (or C++ or whatever) ad nauseum.
It’s as if the author started with a narrative and went searching for the line plots to justify it.
To be clear I don’t doubt the growth of Java amongst new learners is slowing down. Partly perhaps due to an image problem relative to other new languages, but likely also due to having somewhat saturated its “total addressable market” and not repositioning itself in new “markets” in the way that, say python has done.
But it’s still a great language for many categories of problem (e.g. high throughput live service) and has rich ecosystem. Build what you like with whatever tool is good for the job. Stop using bad stats to motivate your recent learning choices and just go build stuff :-)
Enterprises have to be sustainable. If you're furious with Lightbend for making this change to protect their bottom line, then you can always fork one of the last open source versions and maintain/fix that yourself. But most people won't do that, because if they did they'd likely have been active contributors to the project before now, reducing the ongoing maintenance burden on Lightbend and potentially avoiding them feeling the need to close-source in the first place.
Also - anyone who says that x database must be rewritten “because GC” is just making an incredibly un nuanced argument about a nuanced problem. People have built production ready databases in both Java and Go. If you care about low/predictable tail latencies then you have a bunch of other more important problems to solve before you worry about the behaviour of a modern garbage collector. For example: how good is your cache hit ratio? How are your synchronous replication protocols affected by grey failures? That kind of thing.
If your raft state machines are doing IO via some write through cache (which they often are) then having specific machines do specific jobs can increase the cache quality. I.e. your leader node can have a better cache for your write workload, whilst your follower nodes can have better caches for your read workload.
This may lead to higher throughput (yay) but then also leave you vulnerable to significant slow-downs after leader elections (boo).
What makes sense will depend on your use case, but I personally agree with the author that multiple simple raft/paxos groups scheduled across nodes by some workload aware component might be the best of both worlds.
Or if you absolutely need on premise and are small there is the startup program for free enterprise licenses (https://neo4j.com/startups/)
I also agree with the recommendations for "Designing Data-Intensive Applications" and "Database Internals". Though, having read the latter for a book club at $employer, I felt it served better as a sort of "index for the space" for people who already had some DB experience, rather a true introduction.
## Blogs:
- http://muratbuffalo.blogspot.com/
- https://bartoszsypytkowski.com/
- https://decentralizedthoughts.github.io/
- https://www.the-paper-trail.org/
- https://pathelland.substack.com/
## Other web resources
- https://aws.amazon.com/builders-library/ - set of resources from Amazon about building distributed systems
- https://www.youtube.com/playlist?list=PLeKd45zvjcDFUEv_ohr_H... - lecture series from Cambridge
## Books
- https://www.cl.cam.ac.uk/teaching/1213/PrincComm/mfcn.pdf - A great book on the maths of networking (probability, queuing theory etc...)
- if you roll your cluster membership a lot the dotted version vectors which are created by Akka distributed data grow unbounded. Eventually they will start making gossip messages exceed the default maximum size (a few kB IIRC) and fail to send.
- in the presence of heavy GC Akka cluster has a really bad time. Members will flip flop in marking each other unavailable. Eventually this will render the leader unable to perform its duties and you will struggle to (for example) allow a previously downed member to rejoin the cluster.
- orderly actor system shutdown will also fail under high GC, which is problematic as sometimes you need to restart your actor system.
- split-brain resolution is really really hard to get right. The Akka team have recently made theirs open source I believe which is good, but back when we were building with Akka cluster it required a Lightbend subscription.
- If you aren’t all in on Actors, the integration point between Akka and the rest of your codebase can be a little odd. You often feel like you should reach for `Patterns.ask` (a way of sending a message to an actor and then getting a Future back which will complete on a particular response) but then people tell you that’s an Anti pattern.
————
Having said all the above, if you’re able to go all in on the Actor pattern and you’re unlikely to hit high GC then you should give Akka cluster a try. The problems it tackles are genuinely hard and you should build on their hard work if you can. In particular they offer (in distributed-data) the most robust/complete set of CRDTs I’ve yet come across. Many other CRDT libraries expect you to bring your own gossip protocol and transport layer.
On the other hand I feel like projects should try and use open source “building-blocks”, like etcd and zookeeper, when building their distributed systems. Not only does this help iron out correctness bugs, but it also means that more people are familiar with the quirks, limitations, requirements etc.... of these tools. For example, I think I would be frustrated to hear that K8s were implementing their own raft.
- Academics (most often publicly funded via grants and university salaries) do the work for free.
- They are expected to learn to use LaTeX and to typeset their work for free.
- They are expected to copy-edit the papers for free, or else pay a copy editor themselves with, you guessed it, public funds.
- Volunteer Academics (on university time and therefore, again, public money) are expected to review the work for technical accuracy and novelty. If done well this is extremely time consuming.
- Finally, the Journals have the temerity to charge the same universities who produce their product millions of pounds a year in journal subscriptions and Open Access fees.
- Finally finally, none of the Authors are ever paid for their work. Not that it matters, because again: public funding should mean public access.
The most frustrating part is that Academics themselves are locked into this system by the career prospects conferred by prestigious journals/conferences.
I like to hope the ACM and other signatories will face a backlash for this. But they most likely wont
- Academics (most often publicly funded via grants and university salaries) do the work for free.
- They are expected to learn to use LaTeX and to typeset their work for free.
- They are expected to copy-edit the papers for free, or else pay a copy editor themselves with, you guessed it, public funds.
- Volunteer Academics (on university time and therefore, again, public money) are expected to review the work for technical accuracy and novelty. If done well this is extremely time consuming.
- Finally, the Journals have the temerity to charge the same universities who produce their product millions of pounds a year in journal subscriptions and Open Access fees.
- Finally finally, none of the Authors are ever paid for their work. Not that it matters, because again: public funding should mean public access.
The most frustrating part is that Academics themselves are locked into this system by the career prospects conferred by prestigious journals/conferences.
I’m not normally one for beating the “nationalise them” drum, but if there has ever been a case for businesses to be dismantled and put in public hands it’s these parasites.
Sincerely, a Scientist :-)
Without this - or incredibly tightly policed regulation of the existing private agencies, I don't see how an outcome like the one described here isn't mathematically guaranteed?
If you want a simple in memory graph modelling library then check out Apache Tinkerpop. Its great.
Though I'm not sure the US having access to those would make me feel any better :)