But of course, it's complicated.
Throughout of various floating point operations can vary by a factor of 100 on a simple machine. Which is why the standard benchmarks perform a combination of typical calculations. There's the "Gibson Mix" abd "Livermore Loops".
Roy Longbottom is perhaps the most comprehensive compilation of results from testing via these benchmarks.
https://www.roylongbottom.org.uk/Cray%201%20Supercomputer%20...
Essentially, the Raspberry Pi 4 is about 100 times as fast as a Cray 1.
I can find tiny chips that can do specific DSP applications at 160 MFLOPS via fixed-function units, but it's difficult to know if that's comparable.
The FPU on some ARM Cotex-M4+ devices can rival a Cray 1 in general-purpose floating point, and their fixed function units can exceed that 100x. Mostly these are control units for LIDAR devices in automotive applications. They can pull more power than the Apple M1 on my iPad, 20 Watts or so.
How many Cray 1 computers does it take to sit there and listen for me to say "Hey Siri..."?