Lichess gets a big upgrade. It doesn't go as planned
lichess.org
lichess.org
> When everything compiled, I shipped it. And to everyone's surprise, apart from a few bugs I had created while rewriting thousands of lines of code... it worked. It just did. No explosions, no obscure bugs, no memory leak, no performance degradation. That was rather unexpected.
...followed by
> Then we saw the JVM CPU usage rise to alarming heights, with unusual patterns. And no obvious culprit in the thread dumps...
> it was just the JVM that needed some tuning.
Details: https://lichess.org/@/thibault/blog/lichess-on-scala3-help-n...
bit of an editorialized clickbait title really.
Sure the title is jazzed up a lot, but calling it clickbait seems to be a bit harsh.
It is not the articles fault the first sentence was used as the title
INNOCENT dev did THIS, and what happened next will SHOCK you
I clearly don't have the same definition of clickbait.
What does this have to do with clickbait?
it seems there is no universal definition of clickbait. So I based my comment on this:
1. the article was posted on Thibault's own website that doesn't generate revenue per click. Isn't clickbait associated with the possibility of additional ad revenue?
2. Thibault recently called for help on the issue. This post recaps what he went through. He clearly expressed that the code upgrade process didn't go as planned for him. the post and sub-title feels genuine.
3. Clickbait (to me at least) indicates that there's either intent on generating revenue or self-promotion. I fail to see those conditions here
So that's my logic. So why would this be considered clickbait?
Both "big upgrade" and "didn't go as planned" are technically true, but overpromise. A big upgrade could be new features, improved UI, something that affects the user experience.
"Updated Lichess language version, had to tweak cache settings" is a title that isn't trying to sell itself.
Thibault isn't posting low-effort spam for profit, he's pouring years of his life energy into a beloved project that is a gift to the world.
I think showmanship is okay if it's tasteful. The problem with clickbait is that it's manipulative and tacky and dishonest. After you read the article you feel like you've been had.
Having spent several years in the running-JVMs-in-production trenches I would claim the opposite.
> I'm sure there are old-hands who have deployed JVM apps at scale might be unsurprised and unimpressed, but I was surprised.
I've never done anything nontrivial with the JVM, so I'm totally willing to believe I was surprised because of my JVM ignorance, but I've never had to fiddle with my language VMs.
Regardless, I argue there's room for tasteful showmanship without it being clickbait.
Clickbait is trying to bait you into clicking. This definitely falls under that. Is it the worst case of clickbait I've seen? No. But it is definitely clickbait.
The most beneficial title from the reader's perspective is a title which allows the reader to quickly decide whether they need to know the information and informs the reader with the necessary information as efficiently as possible. Hence the now long dead practice of writing in the form of an inverted pyramid (https://en.wikipedia.org/wiki/Inverted_pyramid_(journalism)).
Do you feel this article did something wrong? Because from what I'm hearing it seems like you're saying it wasn't to your taste, which is entirely valid but a different matter.
I don't like click bait. I voiced that opinion. That's all.
Whoever posted it to HN decided to go with the title of the first section of the article.
What's wrong with a title that is a one sentence summary of the article?
So nobody intended clickbait.
Instead, a smug tale of a huge upgrade that went unexpectedly smoothly - with a performance problem swiftly solved. Still a good story, and kudos to the team. Just not the horror show I came for.
edit: Since I got sniped pretty hard for not divulging further, allow me to 'kill the joke' by explaining. This will read more harsh than I intended
Shipping it isn't the problem. In fact, that's the expectation. Just... don't ship it to a place that matters first.
I feel baited by an otherwise cavalier operator, footguns be footguns.
as planned -- I'm not convinced there was a plan
I don’t pay for Lichess, love the UX, and I know it’s a labor of love, but come on. Surely Lichess refactoring aren’t so urgent and user critical that they should just go straight to prod?
Yeah, there should be some kind of shakedown test.
But Nah, not every issue can be tested for within the manpower/$ budget. Issues of load being one of them.
A valid strategy is: send it!, work with the beast a bit, if too unruly roll it back. In this case, they tamed the new beast. Not like Lichess is going to lose billions...
Think about it - clients or requests are really just ephemeral ports...
There are something like 30 thousand of them available by default with Linux
If you use sockets you can have as many connections as your system allows for open file descriptors. The upward limit is...
$ sysctl fs.file-max
fs.file-max = 9223372036854775807
Spawn a ton of concurrent requests to hit the service/API!apachebench is a decent example of the idea; imagine some proprietary sauce with it
It gets cheaper (and in some ways, more effective) if you actually remove the network -- latency goes down
You can also rate limit it in software, but it's really not analogous in some ways (eg: net.core.somaxconn limitations/realities)
Development time is totally understandable though. It's important I leave my green field of comfort
I'm not saying it doesn't take effort, I'm just saying it can and is regularly done
Regression testing with speed. If you test your features as they're added, niche-ness doesn't matter - you know about it and test it
A regression test quickly turns into a load test if you do them all at once
To be clear, I'm not trying to twist the knife. I understand why it wouldn't be done. I just won't pretend it can't
I've been trying to avoid this 'hindsight is 20/20' thing...
I'm not trying to lay blame - just offer thoughts towards improvement, and encouragement against some defeatist thinking patterns
A lot can be done with very little equipment/time investment
It's not particularly expensive or difficult, but it does need investment
> We want types, not boilerplate. Sometimes it's best to let the compiler figure out what things are by itself.
At first I thought type inference was cool too, but more and more I don't like it. I prefer being able to see a variable's type easily, even when skimming over code. Type inference makes code less readable to me.
I wouldn't say "because type inference can be harder to read in some places, it should always be avoided".
A common example of this is a variable that's initialized in one line, fed into another function in the very next line, and then never used again. [1] If the function call that produces it, and the function call that consumes it is named well, the variable's type is often, but not always superfluous.
It's like naming for loop variables. Sometimes, you want a specific, detailed name for them. Other times, you just call it 'i', because what it is... Is not very important. It's a counter, counter goes up until we hit the end of a list, don't think about it too much.
[1] The two function calls could have been mashed together, but that would make the code less readable. The full variable type could have also been included, but its not very relevant to the intent of the code - which is simply[2] feeding the output of one function as an input into the other.
[2] Obviously, when contextual information (like real-world side effects of creating a particular variable - say, a thread, or a record...) that complicates the "Create -> Use" case is relevant, or if the APIs in question have a lot of function overloading, or poorly-named interfaces, or if non-trivial casting is going to take place, it may still be beneficial to explicitly provide the variable type.
EDIT: Also, what do you mean by type ‘coercion’? As far as I know the types in Rust are never silently coerced to anything. They just are what they are.
There are some specific coercion sites and kinds: https://doc.rust-lang.org/reference/type-coercions.html
I'd rather work in a language where nobody has type information. It's a much more even playing field. ;p
Main benefit of static typing is enabling tooling to reason about your code - not using IDEs is throwing a huge chunk of it away. Historically it was easy to get projects that would be too large for real-time tooling - but these days the IDEs got better and you can get 32 or even 64 GB ram into a workstation trivially - I haven't seen a scenario like that in years.
I've also noticed GitHub supports type navigation in some languages, but yeah for nontrivial reviews I'll usually do a checkout anyway.
Non-sarcastically: type systems have their place but are not quite there yet. I say this as someone currently doing Advent of Code in Haskell, which surely cannot be accused of being lacking in type safety. What I would love to see in a type system is the flexibility that (for example) Rails' ActiveRecord adds to a class while still being type-safe. AFAIK, no such system exists (yet) but I would love to be proven wrong sometime in the future.
I think vim-lsp and Neovim's native LSP both have it too
https://stackoverflow.com/questions/66174400/how-to-get-type...
Personally, I agree, Rust is nigh impossible to write without types in your editor; thankfully Rust Analyzer provides them.
var someLongThing <application.someHierarchy.someLongAssPackage.someLongThing> = new application.someHierarchy.someLongAssPackage.someLongThing()
class class1<T extends Comparable<T> >
implements SmallestLargest<T>
T[] values;
This is a spot where inference doesn't work, but has the same problem as the relatively tame example you used.The solution here is not to eliminate the type signatures. The solution is not to wait until compile time to figure out the type of a thing. Nominal typing, as in named types, is the solution here. If you have a 'thing', which also happens to be an application.someHierarchy.someLongAssPackage.someLongThing implements application.someHierarchy.someLongAssPackage.someInterface, anotherLongAssDefinitionForTheReturnTypes but that's an implementation detail, not a type description.
It's an Address, or a customer impression, or a similar product, or a fuel tank. Just fucking call it what it is instead of pretending like you're working on your PhD thesis. Your coworkers aren't robots, or academics, they're human beings who've got real work to do and kissing your architectural astronaut butt/complementing your farts are not on the list. Name things like a human. A wise human, but a human.
let pos_nums = v.iter().map(|s| s.parse::<u64>().unwrap()).filter(|n| n > 0);
Hint: you literally can't in current Rust because closures have Voldemort types, you can't name their type. And even if you could, the resulting type would be horrendous.As it stands, type inference has a tendency to make code reviews more difficult.
Could you share details about how do you do it? I'm very interested!
I find it interesting that you characterise static typing as about safety. I think it's more about communication. (for me this includes "communication to myself when I wrote this seven years ago and didn't think about it since", which I appreciate may be a bit of an edge case). Tests are also a great medium of communication, but a different one - in which case the metaphor becomes swap a sonnet for a sonata? I'd like both please, but maybe not always at the same time :-D
I know you're not saying this, but your wording implies it to a bit... Tests to cover <the things covered by static typing> are only one of the uses for automated testing. There's plenty of places (most of them) where automated testing is just as useful with a statically typed language.
As much as I can get behind val something = new MyClass(), where the type is explicitly mentioned on the line, it's a whole other story for val something = someFunc().
You have no way to know the type of that variable without looking at the function itself which is more work than should be needed.
Just checking quickly a repository on GitHub for example.
Fortunately such bugs are pretty uncommon.
Technically the compiler technically did nothing wrong, but you are still left with a head scratcher because all of the logic seems sound and yet it produces occasional garbage.
I don't know a single language where integers are "promoted" to floats when they get large; certainly no statically typed language. You could argue JavaScript does this, but it's neither compiled nor does it actually have integers. So yes, if a compiler did that it would be a bug.
That said, this is an example of the compiler being wrong.
If a value is changed to a data representation with different comparison semantics, and later compared to a different value of a different representation, it means the original optimization (hand wave over the definition of that) was unsound in the first place. AOT compilers would classify this as a bug that needs to be fixed imminently, and JIT compilers that do this kind of thing use something called dynamic deoptimization to remove the change once its been invalidated.
Scala 3 doesn't do that [1]. The new inference algorithm will, however, prefer a generic method over the same named method without a generic argument in scope when no typing hints are present in the source code when there are two possible matches for an inference completion. This can then cause problems when the generic expects some kind of given parameter:
def foo: String = ...
def foo[A:SomeRequiredThing]:A = ...
//in some other place
foo shouldBe("hello") // foo[String](using SomeRequiredThing[String])
// gets preferentially selected by the scala 3
// inference algorithm, which can lead to surprising
// though not "incorrect" results [also 1]
[1] https://www.youtube.com/watch?v=lMvOykNQ4zsAn IDE can still show the inferred types!
The parameters they are tuning are available in any good VM with a JIT. You won't see these tuning knobs with Go because there's no JIT. These are trade offs. Scala is also known to create more garbage than Java.
The reason for the CPU spikes are code is getting compiled via JIT which takes CPU. If isn't enough space for JIT compiled code then you'll be recompiling code often, burning CPU.
This is similar to database cache tuning. Too little cache and you burn CPU and IOPS. Too much cache and you don't leave enough memory for other operations.
These parameters have good defaults but large apps can surpass them. It's important it's not automatic as this could have dire consequences (oom etc).
This sounds like a bug to me. If you are out of space for compiled code stop compiling. Or at least start slowly raising the threshold at which you recompile something so that the compilation cost is amortized over time.
If your recompiling is using more CPU than was being used before it started something is wrong. You should only be compiling code when it is expected to save CPU long-term.
I wonder if it would be easier for them to just use GraalVM.
Pardon my ignorance. Why can't you set up a performance test in both environments and measure the speed?
If Thibault enjoys Scala 3 and feels more productive, while the overall performance on the same hardware is approximately equal, it's already a win for the project.
Everyone has a performance test environment, not everyone has one that's separate from production. In this case, generating realistic test patterns is probably very difficult, so you get one environment for everything.
> So I forked it and butchered it to remove everything we don't need - which is actually most of the framework.
I guess there is hope that Play framework itself will be migrated to Scala 3 and that the dependency on the fork can be removed, but this is taking on a risk - what if there are security updates to the upstream in the mean time?
The likelihood of getting security fixes in the future is about the same as getting new security holes created by whatever updates. Butchering away everything they didn't need certainly didn't harm security.
Rather than use a big bang deployment. Setup an A/B test, and gather a sampling of data.
I haven’t been able to truly test azul’s claims yet. Java apps that I manage/deploy were on a small scale in regards to requests/sec and I personally didn’t see much difference.
Everything after G1GC was suggested by various helpful experts on HN, Discord, by email and other media.
It seems like a failure of the JVM that someone capable of porting thousands of lines of code with little trouble then fails to tune the JVM. It's like, "we wrote all the code, now the hard part starts, tuning the JVM".
My full rant about it is https://blog.habets.se/2022/08/Java-a-fractal-of-bad-experim...
It's partially my subjective opinion, but it seems that almost every single language design decision that Java made was, in retrospect, a bad one.
Not that I could have done better. It's just that none of it panned out.
But maybe you're saying that the decision was good, but that they failed to actually do it?
Java was the first mainstream garbage-collected language. Not the first GC language, but the first one to get serious traction. It started the post-C++ era.
That was a pretty bold design decision for the time, and one that worked out. The VM was another big one, and it also worked out.
GC wasn't a failed experiment. OOP extremism was. And the latter is a bigger Java standout than GC.
Python predates Java and is still mainstream, and while (mostly) not garbage collected, but reference counted, language design wise it's not a big difference. There have been GCd Python implementations, proving that.
Being successful is not really a language design choice, so I wouldn't count that.
> The VM was another big one, and it also worked out.
As I explain in the blog post, the VM was a colossally bad mistake. The inspiration for the VM sounds like it was better. So looks like they took the ball and started running in the wrong direction.
1. Checked exceptions.
2. Crusade against properties.
3. Bad API is not being actively deprecated. StringBuffer, Vector, Date, Calendar, File. Replacements exist, but old classes are still in core.
Also I think that Java community is crazy with their love of Spring Boot. Like Spring is not magic enough, let's put more magic to autoconfigure its magic. But that's not a failure of Java.
A huge spike in CPU usage would be treated as a regression in other language runtimes and would likely be addressed. The fact that this is solvable with advanced JVM knobs is both good and bad. Good in the way you say. Bad because the complexity of maintaining all those knobs makes it difficult if not impossible to improve the runtime defaults, and every runtime problem becomes a tuning problem.
A spike that applied to all programs certainly would. A spike that applied to only one program? Good luck, IME.
Every release Go gets better and requires 0 changes to the code. I have never needed to fine-tune the GC. I have never needed to spend a month rewriting my code to work with Go 2.0. I have never been nervous to update to a new Go version. I write the code once, and it runs great in prod for years. I love and value these things. I also suspect that Lichess’s use case would perform extremely well in Go out-of-the-box, since it’s just a web app.
I certainly hope the JVM team would like Lichess and other web apps to run well without needing arcane configuration knowledge gathered over years of experience and battle scars in production.
> Every release Go gets better and requires 0 changes to the code. I have never needed to fine-tune the GC. I have never been nervous to update to a new Go version.
I've had the same experience with the JVM, over a longer time period. You hear about this because it's exceptional and interesting, not because it's the norm. "Just a web app" is an incredibly reductive take.
The same problem exists in the Go runtime (it's fundamental to the problem space), and since they don't let you tune it presumably they either guess or hardcode what the value should be; good luck if they get it wrong. Sure, you probably won't hit it with your code - the overwhelming majority of Java users don't hit this kind of problem with their code either.
Err no, Go has GC yes, but most of that tuning seems to be related to JIT/Codepaths etc.
GC will always happen no matter how much heap size you allocate: large heap size will make GC happen less frequently, but depending on the algorithm, it will also increase how much time is needed for each collection. The key point for GC tuning is to keep total GC time and pause time under control with as little memory as possible.
* First, you have to setup monitoring for the garbage collection time: turn on metrics collection and details garbage collection logging.
* Second, tune your total GC time so that it's under 5% or less. Start with a reasonably small max heap size, says 256MB, and keep increasing it if the GC time is still too large. Try to keep the max heap size under 32GB to take advantage of "Compressed OOPs"
* Third, you only need to use more advance flags if you detect large GC pause in your GC logging. Otherwise, you're done.
See https://lichess.org/source for a list of all the services with a more-or-less up to date diagram.
That aside, that thing with opaque types in Scala sounds interesting - but how would the compiler know that the parameter was a UserId and not just a random string?
I took a look and found that it is possible to do something similar here (albeit with more heavy lifting)
https://www.ferreira.io/posts/opaque-branded-types-in-typesc...
I can see the value when writing code in the early stages of a project to both have flexibility and save keystrokes with type inference to help prototype and explore the problem space. Later on, I can see the benefit of switching to typed variable declarations for readability as time goes on and that part of the code base slips away from my working memory.
https://scalacenter.github.io/scalafix/docs/rules/ExplicitRe...
However, from my experience it's actually better to not annotate most things. I annotate methods are that used very often or are "borders" of my API or interface. Otherwise it is usually not needed because it rarely makes the code more clear and IDEs like IntelliJ actually show you the types inline automatically, so you can always see the return type easily.
I tried Scala a long time ago but I didn't like it, it felt every student of Martin Odersky's CS course at the EPFL got to design their own language feature, and they often picked the same one but came up with different implementations.
> At some point I had to rewrite the Glicko2 rating system from Java to Scala 3 [...] No-one noticed broken ratings, so I suppose it worked.
Sheesh. Seems like just the kind of thing for which a couple unit tests would be relatively easy to write and would add a lot of confidence.
But I’m honestly shocked it works as well as it does without tests. I mean I’m not a paying customer and I much prefer Lichess to the other chess website so I guess I have little say…
Does anybody know if Lichess would accept code contributions to add tests?
Scala 3 is also quite different compared to the early times. Unfortunately the reputation is still a bit outdated.
Congrats to lichess and everyone involved who helped!
Brave soul is an understatement.
And this is why it's annoying to use Scala. They just decided to rewrite the syntax of the language
Which languages don't make any changes to syntax over major versions?
If you lose indentation, I want your code to break. It's a feature not a bug!
I'm not a big fan of Python, but significant indentation certainly doesn't seem to have hurt its adoption much...
The decision was directly motivated by Odersky's experience teaching Scala to students for 15+ years, it wasn't made for the sake of fueling drama in the community.
Just use Explicit types. They aren't hard, time consuming, or bad. They are your friends.