Programming language comparison by reimplementing the same transit data app
github.com
github.com
Since I submitted it, though, I posted it to /r/rust and got a lot of feedback. At the time, rust and dotnet were comparable and at the top of the list, with ~10k req/sec. Now rust is far and away the most performant at ~20k req/sec! I also was able to improve Go's performance 30% or so. Still, I want to let the other communities chip in and see if I can improve them.
I was actually just in the midst of exploring how to improve Elixir's results. I'm finding I can almost double my requests per second from simply switching the JSON encoder from Jason to Jiffy. That sort of surprised me since Jason is the de-facto standard, and I thought was super fast.
Also use a release, it will help some things here, especially on "short" benchmark, and make sure that the JIT is run.
Good call on running as a release! I'll try that next.
edit: Ah well, good thoughts but didn't pan out. I updated the logger config to use `compile_time_purge_matching` just to make sure there wouldn't be any logging impact, and I ran the app as a release, but didn't really make any difference in the numbers that I saw.
Personal plug, but have you tried Jsonrs[0]? I typically get much better performance out of it (especially in lean mode, which seems to be the mode you'd want to use for this benchmark) than Jiffy for large JSON encoding workloads.
It's not quite as high as I was seeing with `jiffy` (3,800 req/sec here vs 4,000+ with jiffy), but I'm not confident that was a totally fair comparison. `jiffy` doesn't integrate as nicely with Phoenix, so I was just calling `:jiffy.encode(...)` in the controller and then doing a `text(...)` response. I need to double-check if `json(...)` is doing more work here.
[0] https://github.com/losvedir/transit-lang-cmp/commit/140d693b...
Jsonrs supports encoding protocols by default, but lets you turn it off for a speed boost if you're okay with more Jiffy-like behavior. Plugging it in as Phoenix's JSON library won't turn off that protocol handling, so you're getting the slower operational mode of Jsonrs (which is fine and probably what you want in most cases for correctness, but will definitely eat a few ops/sec in a benchmark)
https://elixirforum.com/t/an-informal-comparison-of-several-...
----
It's not surprising that Elixir is not good at large json handling and complicated data structure manipulation. This reminds me of https://discord.com/blog/using-rust-to-scale-elixir-for-11-m...
At the same time I see that Elixir gave the most consistent response time. Max response time is less than 6x of median compared to ~x10 (Deno / C#) or x23 (Go / Rust) / x58 (Scala), probably thanks to preemtive scheduling of per-request process in Erlang VM
Regarding GenServer, I'm not so sure about that. I suppose I should benchmark it, but intuitively I expect ETS to be better here. There's more overhead in getting the data, but it allows you to concurrently read it from different processes. A GenServer, meanwhile, could respond to a given message faster, but now you're serializing the (e.g.) 50 concurrent virtual users through a single bottleneck. I could have multiple copies of the data in several GenServers, I suppose, at the expense of much more memory use.
These Elixir numbers seem suspect to me, based on my company's extensive benchmarking on many of these same languages. Elixir is somewhat slow at some tasks, but serving up web requests isn't one of them.
> The "billion dollar mistake" is important to me, and while C# has non-nullability sugar in its typesystem (i.e. with ? after a number of types), the type system wasn't as rigorous as I was maybe hoping.
I'm genuinely curious if you have any examples? Nullability checks may be bolted on, but once you enable them, they should consistently prevent you from dealing with any null values that you haven't explicitly allowed. And other than that, the only hole in the type system that I can think of is array covariance, which was unfortunately common at the time (Java had it also) as a way to skimp on generics, but which doesn't occur often today because using generic collections is much more common.
> At one point I had a bug because I did a stopWatch.Elapsed / 1000 by accident instead of stopWatch.ElapsedTicks / 1000. The former is a TimeSpan struct instead of a long like ElapsedTicks, so intuitively it feels like I shouldn't be able to divide it, though it did a best effort and did something to it, though I'm not quite sure what.
This particular one doesn't have anything to do with the type system per se, it's just the way TimeSpan itself works. I'm not sure why "intuitively it feels like I shouldn't be able to divide it", since there are fairly obvious definitions of arithmetic operations on spans and numbers - if T is 10 seconds, then surely T*2 is 20 seconds, and T/2 is 5 seconds? Which is exactly what TimeSpan does by overloading the corresponding operators:
https://learn.microsoft.com/en-us/dotnet/api/system.timespan...
Note that the type of (TimeSpan/double) is still TimeSpan, so the type system is still enforcing proper use - it wouldn't have let you assign it to an int or a double. But it looks like you were just printing it, which is legal for any type:
Console.WriteLine($"loaded stop_times.txt in {watch.ElapsedMilliseconds} ms");I’m not the one you asked, but there’s a difference between you no longer having to deal with null checks and your program no longer having to deal with null checks.
The CLR will still do the null checks and fail your program fast when needed. Other languages will require you to properly handle them (or at least, make it obvious where you chose to bluntly abort the program when they occur)
I also don’t think C# will prevent you from forgetting to do a if null check in cases where you expect values to sometimes be null.
It should throw a warning in situations where a non-nullable value is not clearly non-null when it’s dereferenced. So you won’t get a warning in the body of a method declaration with a non-billable parameter, for example, but you will get a warning if you attempt to pass null into it.
https://learn.microsoft.com/en-us/dotnet/csharp/nullable-ref...
So yes, "will require you to properly handle them" is exactly what I meant.
That approximates it, and likely gives you most of the benefits, but it isn’t the same as strict enforcement of necessary null checks.
There will be cases where the programmer can prove a value isn’t null, but the compiler can’t, and thus, the programmer can choose to ignore the warning.
That’s similar to how Java tries to prove that a variable is initialized when being read. (https://docs.oracle.com/javase/specs/jls/se9/html/jls-16.htm...).
Java makes the safe choice, though. It refuses to compile such programs, thus guaranteeing that every local variable and blank final field will be initialized before first use.
And if you ignore the warnings, that's a deliberate opt-out; why even bother with nullability checks then?
- enable source generation for JSON serialization
https://learn.microsoft.com/en-us/dotnet/standard/serializat...
- use `TryGetValue` to do single dictionary lookup instead of two.
if (TripsIxByRoute.TryGetValue(route, out var tripIxs))
- specify List's capacity, if it's known when creating it, to avoid resizing operations while adding elements to the List var schedules = new List<StopTimeResponse>(stopTimeIxs.Count);The type-magic style of development is so different from what I'm used to, I have periodic crises of conscience where I wonder if I've been doing programming wrong all these years, or if it's an example of a community barking up the wrong tree. This thinking is also what prompts me to look into APL/J/K every so often.
The last time I really tried to get into hardcore FP programming was several years ago with Haskell, but even then I don't recall Yesod (the web framework I tried out, akin to http4s here) being quite so overwhelming.
[0] https://hackage.haskell.org/package/servant
[1] https://hackage.haskell.org/package/snap
[2] https://hackage.haskell.org/package/scotty
Shortly after Elm, I discovered the array languages via J, and felt like I finally found a style of programming that fit my brain properly. Immutable by default, rank-polymorphic, and extremely powerful. Terseness is a feature - when you can write the same program in 1/100 of the code (not an exaggeration), you can explore alternative approaches quickly and find bugs more easily. On the other hand, there's no compiler or fancy type system to find bugs for you, so terseness is also kind of necessary.
Today I'd recommend array language newbies to try BQN or K first. BQN has some features that other array languages surprisingly lack (closures, modules), and only a few idiosyncratic design choices (unlike J which is very weird). K is pragmatic (it has dictionaries and tables!) and is ascii-based, but it has an almost Forth-like minimalism. For someone used to 'import xyz' style programming it's a jarring transition.
It is a NOT a library that you should start out with when you learn Scala, unless you have a mentor.
I think for you, cask was the right choice and it's good that you found it.
When I first learned of FP about 20 years ago I thought it was a terrible idea. But I've convinced myself I have been proven wrong time and time again.
I still think sometime it's complexity for its own sake. Eg. why is it so hard to define a monad? But I enjoyed parts of Scala 2 and later Rust. There are good ideas in FP that lead to programs that are easier to reason about and manage complexity. Languages are borrowing these good ideas and I'm happier for it, algebraic data types, pattern matching and destructuring, result/either, immutability, etc
Scala 2 is still getting regular updates so there's no rush to use 3 just yet, I'd recommend picking the framework first and using the Scala version they support. That makes the development experience much better with access to much easier to use frameworks like play or finatra.
A slightly better starting point for scala 3 + type-safe server building is tapir e.g. https://github.com/softwaremill/tapir/blob/master/examples3/... . With that, you get a declarative definition of your endpoints (+ error types, auth, etc.) that you can use for both servers and clients, which comes very handy when writing integration tests of course.
> absolutely ridiculous the fetishization of extremely complex FP and type-level hacking that goes on in the ecosystem
An alternative way to look at it is that there is a lot of essential domain complexity that gets encoded via the type system to let the compiler do the hard work. That "extremely complex FP" does not arrive out of nowhere - I really recommend at least skimming through the slides from rossabaker, the http4s designer, that motivate where the core type signature comes from https://rossabaker.github.io/boston-http4s/#2
I suppose one of the "features" that I like about the (typelevel) community is that the approach of "worse is better" is not taken, and a lot of effort is expended to make things correct, modular and orthogonal. This has the drawback of increased upfront complexity, that anecdotally pays off the moment your compiler does not error and the program runs as intended.
I am not a scala user but this sounds like a serious violation of composability (ironically), for something so ubiquitous as http. I don't really mind advanced features being available for power users, or later optimizations, but I really dislike being forced to make a mutually exclusive decision between simple and performant, early in the project lifecycle.
I think Rust unfortunately often falls in this category as well with async and multiple http libraries, and even the inofficial doctrine of deferring ownership decisions for later.
In Go there's basically one way to do things, and the few more advanced things you can do can be opted into later. I assume this is a quality inherited from C, and it's extremely pragmatic and creates a composable ecosystem.
But other than that, running a simple hello-world service with http4s really isn't that hard.
object Main extends StreamApp[IO] {
val helloWorldService = HttpService[IO] {
case GET -> Root / "hello" / name => Ok(s"Hello, $name.")
}
override def stream(args: List[String], requestShutdown: IO[Unit]) =
BlazeBuilder[IO]
.bindHttp(8080, "localhost")
.mountService(helloWorldService, "/")
.serve
}
Yeah, there are a couple of things that are more complex than with other simple webservers. For instance, what is the "IO" thing doing there? Well, http4s flexibly allows you to choose different libraries for handling concurrency without even knowing those libraries. Very few other programming languages are even powerful enough to support something like that, so it obviously confuses people.And then you need a StreamApp (there are other alternatives but that's the one from the documentation) and people might wonder why - they want a simple request/response webservice and nothing "streaming".
Other than that, I don't think the code is overly complicated.
But it turns out there's an alrernative. Tapir is emerging as that "pure FP, but approachable" HTTP framework. It allows developers to work at a higher level, turning the HTTP backend into an implementation detail. Simply define the application using Tapir, then choose the better backend for the use case.
(BurntSushi is Andrew Gallant, creator of ripgrep and other excellent Rust crates, and writes code that is especially good IMHO for learning how to write real-world Rust libraries and programs.)
https://github.com/gothinkster/realworld
Not as performance-focused with benchmarks, but a good point of comparison for various languages and frameworks implementing common behavior.
Looking forward to the reports for plain Java and Nim.
I would need to profile the code, but the startup time being bad for Deno seems like maybe a combination of the code in here being unoptimized:
https://github.com/denoland/deno_std/blob/0ce558fec1a1beeda3...
(Ex. Lots of temporaries)
And usage of the readFileSync+TextDecoder API instead of readTextFile (which is also a docs issue since it suggests the first one). It seems the code loads the 100MB into memory, then converts to another 100MB of utf8, then parses with that inefficient csv parser. The rust and go versions look to be doing stream/incremental processing instead.
I’m 99.9% sure Deno doesn’t handle multiple cores at all. Everything in the node ecosystem is one big asynchronous single-threaded event loop.
That's right because Google's Dart (2010's) was heavily influenced by Microsoft's C# (2000's) which was heavily influenced by Sun's Java (1990's)
I've often wondered if would have been better if the companies had worked together to produce a standard VM language
In Go you want to start with the package examples. CSV has NewReader(). I had the same experience with Java... you know what a function takes, but where do you get one of those?
Still I would try a framework with Go, it seems odd to not use one when a framework is used with the other languages.
But that’s a fair take. Go’s web server ecosystem has diversified a lot too, and I’m sure his pick would’ve affected the outcome a reasonable amount if he had chosen fasthttp or gin for comparison.
Node really craps the bed when you start pulling in libraries, though. And there are lots of performance surprises such as fetch being quite slow vs the original http client, etc.
> This repository implements the same simple backend API in a variety of languages. It's just a personal project of mine to get a feel for the languages, and shouldn't be taken too seriously.