It has to do with offloading computation to an external entity.
For example, lets take a simple example: CPU + DSP. I have an Atari Falcon here which sports something like that, but quite a few modern SoCs also have DSPs included (QCOM, TI, ...), or you have SBCs (single board computers, like Raspberry Pi) which sport an additional DSP chip on it.
Now what do you think the process will be to do the computations on the DSP?
You will need at the very minimum some way to hand data over to the DSP, execute some code, and retrieve whatever result back to the CPU. This can be as simple as just handing around memory pointers (in SoCs with "unified memory"), or as complicated as specifically triggering the communication line between these two chips (some fast serial connection, ... whatever). That also means you need buffers for that, i.e. the CPU has to somehow provide buffers to some API that deals with the handover; at the very minimum you need a buffer for the input data, the actual code, and for the output (unless you reuse the input). Now, you define some code to do whatever computation, who figures out the sizes of these buffers? For the code itself it can be easy, depending, but input and output?
From there it gets easily more complicated, and that is the very reason why there is so much "boilerplate" in APIs dealing with GPUs, and some drivers actually support OpenCL, but I'm not sure right now if that is made available in browsers even. This aims directly at "compute", but also brings all the aspects you need for properly offloading that with it. There is simply no way around that.