Energy Efficiency across Programming Languages (2017) [pdf]
greenlab.di.uminho.pt
greenlab.di.uminho.pt
If your company's business is really in running some algorithms similar to binary search or the n-body problem, you should probably not be checking which language to choose that is most energy efficient, but which library or framework suits you the best. E.g you can perfectly well use Python + Pandas or NumPy even though everybody will agree that Python in itself is terribly inefficient. If you really want to go to the extreme in saving energy you should go all the way and code your algorithms in assembly. But it's obvious why this is probably not a very wise choice.
Everybody who knows a bit about IT will understand this, but the danger is that some project managers will read one of these papers and draw the conclusion that the team should switch to C or Rust because it is more sustainable ..
IMO, the first wave of optimizing software for energy consumption has already happened via virtualization. Ideally, I can run the same workload with less metal/watts.
I feel like the next round of optimization will be more problematic. You could re-write software to take advantage of hardware acceleration features that are baked into the silicon to ideally speed things up/reduce power consumption, however this comes with its own issues. You need to be very familiar with your hardware. Integrating much more tightly means your married to your metal... kinda like CUDA lock-in in the GPU ecosystem.
I don’t know how true it is in the HPC space, but the more optimization work I do the more I realise how horribly inefficient 99% of programs are. Most devs these days just don’t care about making software fast, and when they do they have no idea where to start or what will help.
Last year I spent some time optimising CRDTs (used in collaborative text editing). Automerge (a well known library) took 5 minutes to run a certain benchmark. Yjs (considered a mature, well performing library) took 1 second to run the same test. After spending some time optimizing, I have code now which can do the same work in 4 milliseconds. I’m not using any special hardware - this is just in straight forward rust.
The remarkable part is that automerge isn’t written in an uncharacteristically bad way. Sure, it uses immutablejs - but plenty of modern javascript programs do something very similar. Even yjs, despite being the target of some careful optimization work, is still 200x slower than it could be. - and thus, wasting 99.5% of its energy budget.
If even yjs can be sped up by 200x, how much do you think most enterprise software could be sped up, if people tried? 500x? 10000x?
I don’t think we need deeper integration with hardware. We just need more software developers who care at all about performance.
Virtualisation wipes out 20%+ of the performance of the machine, depending on what you’re doing. But we pay that cost gladly because it’s marginally easier than sharing a Linux kernel between multiple processes.
I think if our CPUs had capped out 50x slower than they are today, the average user would hardly notice any difference. The only difference would be that software would be more expensive to make, because software engineers would need to actually understand how the computer works to make software perform. Things like the virtual DOM wouldn’t be fast enough to use at scale.
Video games and zoom wouldn’t be quite as pretty. But we’d live.
Like 10x most? You compare something that is basically one hot loop with something that at most spends 1% of its CPU time in hot loops, the other being just random code paths and calls to other, often optimized libs (e.g. for making network connections, db communication, etc)
A flat performance profile isn’t a sign everything is fast. It’s usually a sign everything is uniformly sluggish. And yeah, that makes optimization harder. But it’s far from impossible.
Take rendering html for example. I’ve seen rendering (just the rendering part, not database lookups) take over 200ms in nodejs. Generating a string with the equivalent rust compile-time templating engine can spit out the same html in microseconds.
There is an insane amount of performance on the table with modern CPUs. Just about everything short of network round trips and stable diffusion and should complete instantly.
I'm sorry but this is far from the truth.
The absolute majority of software written today is not optimized. There has been a huge shift away from efficient close-to-the-metal programming that was popular in the 90s and earlier, to what we have today. There is enormous energy and efficiency benefits we could reap today, without "problematic" optimizations. We just need to start optimizing at all.
If it's not problematic, then why is it not being done already? Obviously, it's the economics... which are problematic.
When it comes to other languages I don’t think it’s that simple either. Yes, these benchmarks are not representative of what most people do at work, but I’ve seen orders of magnitude of difference between dynamic languages like Ruby or Python vs languages like Rust or Go when handling web traffic. Often web apps spend 50-70% of the time to handle request doing actual CPU work
Also, for performance critical sections, you can bypass the GC by importing "C" and "Unsafe" to get manual memory management and pointer arithmetic.
https://dgraph.io/blog/post/manual-memory-management-golang-...
FFI is available in basically every single language, C# was actually very explicitly made for this use case.
And Java originally was literally made for embedded devices :D it is still running on every single SIM card and bank card, but there are many other solutions targeting embedded, for example microEJ.
As mentioned you can also bypass the GC for performance critical sections.
Benchmarks Game is a fun exercise to kill some time on, but it doesn't indicate anything about anything. The solutions across different languages don't even have the same algorithmic complexity. How can these possibly be compared?
Here's another, for the server use case:
https://www.techempower.com/benchmarks/#section=data-r21&tes...
I am stunned that a JS tool came out on top.
> When a measure becomes a target, it ceases to be a good measure.
If you're deciding what language to use for your next project, there are better ways. Ask yourself these questions.
- Can developers be hired/trained to use it?
- Will they be happy using it?
- What will our development velocity be on this language? What will be the operational burden of keeping it running in production?
- Do we need to rely on third party code? Does that code already exist and what is it's quality?
Importantly, how will the answers to these questions change over time? Will this language be just as easy to hire for in 5 years time?
Comparing languages by benchmarks is a waste of time.
All those are subject to physical constraints though. Much of the worlds code is embedded, and thus severely constrained in both compute and power. That limits your choices, but within those choices your criteria are still absolutely the correct ones.
I wanted to know why, and I found this article by the author:
https://just.billywhizz.io/blog/on-javascript-performance-01...
also: https://github.com/just-js/just/issues/5#issuecomment-778673...
leetcode even has many versions of each problem written by different people so you could see e.g. the distribution of expected runtimes for JavaScript programmers Vs Python programmers or whatever.
It wouldn't quite be fair on the language. E.g. I would expect Typescript to give better performance than JavaScript simply because the fact that you're using Typescript means you know what you're doing more than a JavaScript programmer.
Would be cool to see. Unfortunately I couldn't figure out a way to get unbiased sample programs for each problem.
https://haslab.github.io/SAFER/scp21.pdf
"In addition, we further validate our results and rankings against implementations from a chrestomathy program repository, Rosetta Code., by reproducing our methodology and benchmarking system."
Still the overall results mostly match my intuition.
With exploratory data analysis it is important to state how outliers will be identified and treated.
https://www.itl.nist.gov/div898/handbook/prc/section1/prc16....
The data tables published with that 2017 paper, show a 15x difference between the measured times of the selected JS and TS fannkuch-redux programs. That should explain the TS and JS average Time difference.
There's an order of magnitude difference between the times of the selected C and C++ programs, for one thing — regex-redux. That should explain the C and C++ average Time difference.
Even without looking for cause, they seem like outliers which could have been excluded.
Without those outliers you'd be telling me that "the overall results mostly match my intuition".
> Through also measuring the execution time and peak memory usage, we were able to relate both to energy to understand not only how memory usage affects energy con- sumption, but also how time and energy relate. This allowed us to understand if a faster language is always the most energy efficient. As we saw, this is not always the case.
> this is not always the case
Just almost all the time, with a few exceptions. Execution time and energy consumption are very strongly correlated.
Time taken is the single biggest factor in how much power is consumed in modern CPUs. The idea is to complete any task quickly and put the idle cores to sleep to conserve power.
> Execution time behaves differently when compared to en- ergy efficiency. The results for the 3 benchmarks presented in Table 3 (and the remainder shown in the appendix) show several scenarios where a certain language energy consump- tion rank differs from the execution time rank (as the arrows in the first column indicate). In the fasta benchmark, for example, the Fortran language is second most energy effi- cient, while dropping 6 positions when it comes to execution time. Moreover, by observing the Ratio values in Figures 1 to 3 (and the remainder in the appendix under Results - C. En- ergy and Time Graphs), we clearly see a substantial variation between languages. This means that the average power is not constant, which further strengthens the previous point.
There is a strong correlation, yes, which makes sense because more time will likely correlate to more work done by the hardware. However the paper proves that there are exceptions to this rule. That's interesting given that most people, including you, would have flatly assumed that performance is a proxy for energy efficiency.
You could argue that it's not a finding which is going to be relevant to most software engineers, but it might be interesting if you are designing or implementing a programming language for instance.
There are a few exceptions, but no mechanism that explains that from the paper. So statistical noise.
If anyone is designing a language to be energy efficient, they would do what is obvious - do less work by generating more efficient code (like you said) or make it easy to parallelise work so cores can go back to sleep quicker (like I said). They’re not going to design for something statistical anomalies.
They have a meaningful new finding, which may provide the missing piece to some solution later on.
If we disregarded all scientific inquiry which did not result in an immediate solution to an existing problem, computers probably would not even have been invented by now.
Eager to hear your thoughts on this once you’ve organised them.
You've been just wrong enough on all points that I could almost believe it's your paper and you want to bait people into pulling out the best bits and quoting them lol.
"Eschew flamebait. Avoid generic tangents."
"Don't be snarky."
Clearly not true. This is like when people say "IQ tests don't measure anything" but what they mean is "IQ tests don't measure what I think you think it does"
You assume we are all incapable of understanding the nuance of the Benchmark game, maybe some of us are. Not all of us!
>The solutions across different languages don't even have the same algorithmic complexity. How can these possibly be compared?
You can compare anything you want. If you feel a particular languages solution is sub-optimal, you may go fix it rather than complain. There is no perfect way to compare language performances, this way is reasonable. If you know of a better way, set it up and do it.
I’m not that interested in debating whether gamed benchmarks measure anything useful. You’re saying they do. Let’s agree to disagree.
Obviously not.
Seems like the combination of climate emergency and my programming language beats your programming language has a fascination.
At-least the OP posted the conference paper rather than a jpg of one table out-of-context.
From looking at the code snippets, a big issue with this study becomes clear - it doesn't reflect how languages like Python are used in practice.
In practice, the "hot loops" of Python are in c/c++/fortran/cython/numba/... i.e. Python code usually makes use of a vast ecosystem of optimized science/maths/data science libraries. Whereas the study code is mainly using pure Python.
It's an issue with the methodology that the programs are written specifically for this study; creating an artificial situation.
If you use python as glue, it should be compared to other languages used as glue.
This is a vast overgeneralization. You'll find these optimized loops in optimized packages, done by engineers with enough experience. In practice, there is a lot of slow running code. And that includes "optimized" code that runs on numpy or similar but simply doesn't lend itself very well to being optimized in that way.
Python is convenient for short scripts, but very slow to execute.
Additionally, going through the FFI from C++ to Python and back to C++ has a cost. Eliminating this overhead was one of the motivations behind Julia.
Not correct — the programs were taken from the benchmarks game.
How would this be comparing apples to apples? The paper’s main objective is to compare programming languages in their own native calls. It is like properly translating the Bible or Shakespeare to other languages: some translations will be more textual than others, loosing all the nuances, specially if compared to the original transcript.
If you have python glue calls to other processes written in other languages, then you would be creating noisy in results. It is not the purpose of the paper at all, and it is outside of its scope.
Also, the often claimed negative of Java, memory overhead is relevant here — a GCd language operates best with a deliberate overhead over the strictly necessary memory. Java’s GCs are quite “lazy” in that they will not collect unused objects until it is deemed necessary, which is in line with energy efficiency.
I don't understand much about the theory and practice of implementing a GC'ed language, but I figure the problem is "locking" such an embedded object, because it doesn't have an obvious reference. One could take its enclosing object's reference but I suppose that leads to complications. What about functions that need a pointer to an embedded object in order to read or write to it? etc.
I've had only short encounters with Java, but it (OpenJDK) did bite me in 2016, when I had to process millions of entities. I'd often run into low-memory situations and my machine would completely lock up for a minute or so, before the GC would give up and terminate the process. There was no nice way to solve the issue, other than allocating all my entities SOA style. This helped with the GC problem since it reduced my GC object count from #entities to #columns. But it isn't nice to read/write the code like that, nor is it good for the cache splitting up every little integer.
I hear that GC's have improved much recently, removing the performance issues almost entirely. However, I believe the problem remains that a lot of code is harder to write because you can't simply pass a pointer + length to a function or do a memcpy. I figure this is one reason why there is so much abstraction in Java, with proxy objects and virtual methods and the like.
If the "embedding" problem did not exist, there might not even be such a great need for GC rocket science, because there was a lot less need to have millions of objects in the first place.
I’m half-joking, unfortunately.
The current plan for Java is to divide the objects into 3 buckets, one being the current ones, they are distinct from the others in having identity.
The second group looses their identity, but they will retain their nullability. A great example for that would be the java.time package with e.g. LocalDate. Here two instance with the same fields will be semantically equal and replaceable, but there is no meaningful zero value as there is with numbers. If they would auto-initialize to 1970 that would probably cause some serious, hard to debug bugs later on. So they retain their nullability and they are tear-free. At initialization time they will only get visible in a state that is meaningful.
The last group will be similar to current primitives. They have a designated zero value, and seeing them in a consistent state is not guaranteed (in java, changing an int value even under race conditions is free from introducing values out of thin air, that is only values can appear that were actually set from any thread). In case of a complex number that you are about to change in two steps, another thread is free to observe it in an inconsistent state.
I really like this plan, and this way of thinking of object semantics actually helped me in other PLs as well. Do note that they don’t give any semantics on where the objects get stored, this is up to the VM implementation. In practice what it allows for is copying/flattening/stack-allocation of the second bucket, possibly with some clever encoding of null values a la Rust’s optionals. And the third group will allow for implementing custom numerics with basically no overhead.
[1] https://en.bitcoin.it/wiki/Script
Energy Efficiency across Programming Languages [pdf] - https://news.ycombinator.com/item?id=24642134 - Sept 2020 (158 comments)
Energy Efficiency Across Programming Languages - https://news.ycombinator.com/item?id=21950341 - Jan 2020 (1 comment)
Energy Efficiency Across Programming Languages (2017) [pdf] - https://news.ycombinator.com/item?id=19618699 - April 2019 (1 comment)
Energy Efficiency Across Programming Languages - https://news.ycombinator.com/item?id=15249289 - Sept 2017 (139 comments)
Picking a programming language has virtually no bearing on this. There are certainly some languages that can run faster in some scenarios, but these microbenchmarks have little relevance to any practical reality at scale.
For me, I like to go though hell instead of around it. Don't try to make one server sip the power obsessively. Make that one server do as much as possible, since you are already paying a fixed, idle power cost.
I think picking things like SQLite vs SQL Server have a much bigger impact on power consumption. These are less disruptive choices than swapping programming languages as well.
How so? Usage of, e.g. RAM and memory bandwidth shows huge variation depending on which programming languages are chosen - and that tends to be the most relevant bottleneck wrt. aggregating workloads onto a single server, or a limited number thereof.
If your PL is extremely memory hungry, or doesn't parallelize efficiently, you're going to bottleneck a lot sooner.
I honestly wonder how much time and energy capacity could be freed up world wide if we just used tools optimized for hardware and network performance and efficiency.
Therefore, the real question becomes: Which language will reduce my compute budget most?
Bonus: here's "Ranking Programming Languages by Energy Efficiency", from 2021 [1]
[0]: https://twitter.com/Czaki_PL/status/1569636020475265025
These benchmarks may not be representative for PHP as the usual way of operation for it is not some hot loop numerical compute done continuously.
Remember Jevons' Paradox: Efficiency leads to higher consumption, because that which is efficient will be used more, which in turn leads to higher consumption of the fuel by which its efficient use is measured.
Haven't looked closely at the other problems, but it's apparent to me that the solutions are not even trying to be similar, so comparing their efficiency is near useless.
the problem in question was the k-nucleotide one, IIRC:
https://github.com/greensoftwarelab/Energy-Languages/blob/13...