Thinking Machines – Introduction to Data Parallel Supercomputing (1989)
youtube.com
youtube.com
Which was enough to really have CUDA-like or OpenCL-like code back in the day. Operations could be compiled into 1-bit commands to be executed in parallel across 4096 "SIMD-lanes / threads" (much akin to CUDA-threads).
A lot of the research from the CM2 / Thinking Machines was translated into modern GPU code (parallel prefix sum, radix sort, etc. etc.). The research done back then really lays the foundation upon today's embarassingly parallel works.
There were some other computers that came before CM2 of course. But CM2 + Thinking Machines is very clearly part of the great history of SIMD-compute / PRAM model of compute / Parallel vectors / etc. etc.
The CM2 had (well, those delivered to customers, anyway) a minimum of 8,192 processors, and up to a maximum 65,535. I worked on one for a number of years. I heard (but never saw) that for internal development at Connection Machines each developer had a single 512-processor board to work on.
You're spot on about the single-bit processor thing though. In their *Lisp implementation, you could declare an integer to be as many bits as you wanted, up to the max 128k bits each processor had.
Today's compilers that generate a similar set of code are OpenCL, CUDA, DirectX / HLSL, Opengl's GLSL, Apple Metal, AMD HIP, and Intel ISPC.
-----------
Today's computers (or really, GPUs), aren't 4096x wide devices or 65536x wide devices. GPUs are 32-wide or 64-wide natively, and then MIMD'd into parallel parts after that. (AKA: CUDA has the Grid -> Block -> Thread model. I'm pretty sure that Star-Lisp was only Grid -> Thread, with no "intermediate" block in between).
----------
You probably can get a similar effect as the original CM-2 by compiling into AND/OR/XOR and shift-instructions for AVX512 though?
Or maybe AND/OR/XOR + shift instructions on AMD's CDNA2+ processor (64x wide and 64-bit, for 4096-wide SIMD per core). It'd be pretty terrible, all else considered, because modern GPUs have a slight MIMD factor there.
Lots of other stuff here http://www.softwarepreservation.org/projects/LISP/parallel#C...
Code in other languages still used the FPUs, if available -- but they paid the "transpose in, transpose out" overhead on every operation.
From at least ~1989 on there were a bunch of machines of various sizes and states of assembly around the building at Thinking Machines. You'd connect to the appropriate front-end and cmattach whatever geometry needed. Certain groups did have dedicated CM-2s, particularly those needing specialized configurations (framebuffer, DataVault, etc).
If you knew where the machine running your code lived, you could go sit late at night and watch the LEDs throbbing. The machine really did have a presence.
We’ll get there eventually.
https://longnow.org/essays/richard-feynman-connection-machin...
[1] https://mission-base-creations.myspreadshop.com/ [2] https://www.mission-base.com/tamiko/cm/cm-tshirt.html