Which Programming Languages Use the Least Electricity?
thenewstack.io
thenewstack.io
So if you have a fixed deadline, a 2x efficient language can save you 4x the power.
intuitively i agree but how does this fall in line with the 'race to idle' concept?
would love to see some graphs with operations-per-watt for different cpu/arch.
So in this world more “efficient” task can be run on slower lower powered cores. Stuff that needs to happen as quick as possible is don’t on faster, higher powered cores.
When you can, you just turn off the bigger cores completely.
I don’t have benchmarks, but there must be some good power savings in it if ARM have built their BIG.little architecture around it, and Apple followed suit with their processors.
Also, I am not sure it is correct that energy / op is constant, even at a given voltage? I'm no expert but I read about this when I was thinking about overclocking my device. At a given voltage, higher frequencies still use more energy due to the capacitance of all the components on the device. This is why overclocked energy dissipation scales super-linearly with voltage. Because you aren't just increasing the voltage you are increasing the frequency as well.
Approximately all computer processors.
For computer processors, DVFS isn't used like OP suggests, it's really more like Dynamic Temperature And Power Scaling (DTAPS), because it's meant to extract maximum performance from available power and thermal budget. Also while CPUs _do_ change power and frequency at >1 kHz, they're intentionally designed _not to_ downclock in the example GP gave (inter-frame idle).
That is not true. For any given node process, with the same CPU design there is an optimal energy efficiency curve with regards to its clock speed. Any lower doesn't save you energy per workload, higher means you are paying Exponentially more energy.
exchanging forth messages between cores seems so high level
I wonder if Moore is still thinking about improving these
This is because they used the source from the language benchmark game, which measures some combination of how fast a language is, and how much effort its fans are willing to put in to micro-optimise the programs for the benchmark. In the case Javascript developers have clearly spent more time on it than Typescript developers.
What I find odd is that the table shows Typescript as 7x slower than Javascript. That can't be measuring the same thing, so not sure what is going on there.
Including type checking and transpilation in the timing, I assume.
fannkuchredux.go
https://github.com/greensoftwarelab/Energy-Languages/blob/5c...
fannkuchredux.java
https://github.com/greensoftwarelab/Energy-Languages/blob/5c...
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
I converted the ones I could, like:
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
:but I couldn't get the spectral-norm program to type check.
// The Computer Language Benchmarks Game
// https://salsa.debian.org/benchmarksgame-team/benchmarksgame/
//
// contributed by Ian Osgood
// Optimized by Roy Williams
// modified for Node.js by Isaac Gouy
// multi thread by Andrey Filatkin
import { Worker as NodeWorker, isMainThread, parentPort, workerData } from 'worker_threads';
import * as os from 'os';
enum MessageVariant {
Sab,
Au,
Atu,
Exit,
}
interface SabMessage {
variant: MessageVariant.Sab;
data: Float64Array;
}
interface AuMessage {
variant: MessageVariant.Au;
vec1: UVWField,
vec2: UVWField,
}
interface AtuMessage {
variant: MessageVariant.Atu;
vec1: UVWField,
vec2: UVWField,
}
interface ExitMessage {
variant: MessageVariant.Exit;
}
type Message = SabMessage | AuMessage | AtuMessage | ExitMessage;
interface UVW {
u: Float64Array;
v: Float64Array;
w: Float64Array;
}
type UVWField = keyof UVW;
const bytesPerFloat = Float64Array.BYTES_PER_ELEMENT;
if (isMainThread) {
mainThread(+process.argv[2]);
} else {
workerThread(workerData);
}
async function mainThread(n: number) {
const sab = new SharedArrayBuffer(3 * bytesPerFloat * n);
const u = new Float64Array(sab, 0, n).fill(1);
const v = new Float64Array(sab, bytesPerFloat * n, n);
const workers = new Set<NodeWorker>();
startWorkers();
for (let i = 0; i < 10; i++) {
await atAu('u', 'v', 'w');
await atAu('v', 'u', 'w');
}
stopWorkers();
let vBv = 0;
let vv = 0;
for (let i = 0; i < n; i++) {
vBv += u[i] * v[i];
vv += v[i] * v[i];
}
const result = Math.sqrt(vBv / vv);
console.log(result.toFixed(9));
async function atAu(u: UVWField, v: UVWField, w: UVWField) {
await work({ variant: MessageVariant.Au, vec1: u, vec2: w });
await work({ variant: MessageVariant.Atu, vec1: w, vec2: v });
}
function startWorkers() {
const cpus = os.cpus().length;
const chunk = Math.ceil(n / cpus);
for (let i = 0; i < cpus; i++) {
const start = i * chunk;
let end = start + chunk;
if (end > n) {
end = n;
}
const worker = new NodeWorker(__filename, {workerData: {n, start, end}});
worker.postMessage({ variant: MessageVariant.Sab, data: sab });
workers.add(worker);
}
}
function work(message: Message) {
return new Promise(resolve => {
let wait = 0;
workers.forEach(worker => {
worker.postMessage(message);
worker.once('message', () => {
wait--;
if (wait === 0) {
resolve();
}
});
wait++;
});
});
}
function stopWorkers() {
workers.forEach(worker => worker.postMessage({ variant: MessageVariant.Exit }));
}
}
function workerThread({n, start, end}: {n: number, start: number, end: number}) {
let data: UVW | undefined = undefined;
if (parentPort === null) {
return;
}
parentPort.on('message', (message: Message) => {
switch (message.variant) {
case MessageVariant.Sab:
data = {
u: new Float64Array(message.data, 0, n),
v: new Float64Array(message.data, bytesPerFloat * n, n),
w: new Float64Array(message.data, 2 * bytesPerFloat * n, n),
};
break;
case MessageVariant.Au:
if (data === undefined) {
throw Error('Au received before Sab');
}
au(data[message.vec1], data[message.vec2]);
parentPort!.postMessage({});
break;
case MessageVariant.Atu:
if (data === undefined) {
throw Error('Atu received before Sab');
}
atu(data[message.vec1], data[message.vec2]);
parentPort!.postMessage({});
break;
case MessageVariant.Exit:
process.exit();
}
});
function au(u: Float64Array, v: Float64Array) {
for (let i = start; i < end; i++) {
let t = 0;
for (let j = 0; j < n; j++) {
t += u[j] / a(i, j);
}
v[i] = t;
}
}
function atu(u: Float64Array, v: Float64Array) {
for (let i = start; i < end; i++) {
let t = 0;
for (let j = 0; j < n; j++) {
t += u[j] / a(j, i);
}
v[i] = t;
}
}
function a(i: number, j: number) {
return ((i + j) * (i + j + 1) >>> 1) + i + 1;
}
}The performance of the same program written in TypeScript and JavaScript should always be identical. TypeScript doesn’t add anything to the runtime.
If they got this simple thing wrong, what else did they get wrong?
Here’s the flawed study so you can see for yourself.
But that is only testing the run time of a program. What this paper does not cover, is for example the energy needed to compile the program and of course, the energy needed to write and debug the program. The whole reason not to write "everything" in C is, that there is a real-life tradeoff between having the fastest possible program which also would be the most efficient, and the effort to create it. Higher level and dynamic languages have been created to create correct programs with less effort. Mostly of expensive programmers, but development time obviously goes along with energy usage for the development machine, compiler runs, debugging effort.
But, you even need to go to C if you would like to achieve that, e.g., using rust gives you nice level of abstractions and is highly efficient.
Maybe their benchmarks were mostly single-threaded. But in general this is the first thing to look at.
What I was trying to get at is that if two languages perform the same task, but one of them uses multi-threading to accomplish it, "wall-clock time" is extremely likely to be a misleading proxy for power consumption.
It boils down to the fact that a CPU using all of its cores uses more power than a CPU using only one of its cores. This is easily checkable by listening to the fans on a fully loaded desktop PC, or by using software like "Core Temp" which can display CPU power consumption on some CPU models.
This is true if your cores are all the same and running at the same frequency, but voltage/frequency scaling make things more complicated and heterogenous cores like in essentially every modern ARM processor completely destroy it. If you have a bunch of slow cores that are half as fast but use a quarter of the power [1], using threads to split up the work so you can still meet your deadline can definitely save power.
1: https://www.semanticscholar.org/paper/Power-aware-task-sched...
"C Is Not a Low-level Language" https://queue.acm.org/detail.cfm?id=3212479
> On a modern high-end core, the register rename engine is one of the largest consumers of die area and power. To make matters worse, it cannot be turned off or power gated while any instructions are running, which makes it inconvenient in a dark silicon era when transistors are cheap but powered transistors are an expensive resource.
OTOH there could be a Jenkins paradox where faster compilation results in more compiler consumption (e.g., iterative development). As you say (and I see a comment above addressing this) what are we optimizing for?
Of course this won't work for every program. And with current tech you're unlikely to drive FB with a warehouse full of arduinos ;)
First DDG hit, only skimmed it, but it seems to cover the idea pretty well: https://circuitdigest.com/microcontroller-projects/arduino-s...
For my Tasmota based devices, increasing the sleep time in the main loop to 250ms decreases power draw by 40%. They now might miss button presses (seems Tasmota polls?), but that's a non-issue for pure actors.
My home server can WoL on unicast packets, so that could be used as an "interrupt" to wake the machine from standby. But then you need a suitable workload that allows for substantial sleep time (e.g. wake up for 3s every 30s). Or you could schedule minutes precision polling by waking up via RTC.
Saving power when serving even a few https requests peer second with a sub 10ms response time - as I said, forget about it, at least with x86 hardware as we have today.
That site has build info.
There are probably more options they missed for the other languages.
There are more features on top of this but many require the GC or other memory allocation, which I would say disqualifies them from the most limited of embedded contexts.
In fact in the updated version with Julia added[1] , Julia takes very long to do anything since the Julia warmup is so awful. But somehow even though Julia takes 2.57 times longer to compute results (357%) than Rust and 93x more memory for fannkuch-redux, it's using only 23% more energy. All the while doing tracing JIT and rewriting memory segments to set for execution.
[1] https://sites.google.com/view/energy-efficiency-languages/up...