IBM 801 deserves to be called the first RISC machine, because its design methodology had the explicit goal of achieving a greater performance than IBM 370, by simplifying its instruction set, and it introduced all the principles later endorsed in the Berkeley and Stanford RISC designs (which later evolved into SPARC and MIPS).
CDC 6600 was not created by the simplification of an earlier architecture. It was more complex than the previous very simple CDC computers.
Nonetheless, it was designed to achieve the maximum performance permitted by the available technology and both James E. Thornton and Seymour Cray were extremely competent computer designers, so many of their design decisions coincide with those that were also preferred many years later for the RISC CPUs.
It should be noted that CDC 6600 had a simple instruction set relying on fast register-to-register operations, but nonetheless its ISA was not too simple, as in some misguided RISC designs. For example, it was one of the first, if not the first ISA which included indexed addressing with auto-update of the address register, which are very useful for implementing maximum-performance loops that access arrays (instruction pair fusion is a greatly inferior solution to having 1 bit per load/store instruction specifying auto-update).
IBM 801 also had addressing modes with auto-update, which were inherited from it by ARM, HP PA-RISC and IBM POWER. Aarch64 has also inherited them from 32-bit ARM. While x86-64 has addressing with auto-update only in a few special instructions, it has an alternative method for achieving the same performance in most cases, by having indexed addressing with up to 3 components and with scaled indices (taken from DEC VAX). This allows the use of the loop counter as also the index register for accessing multiple arrays, which eliminates the need for separate index updating instructions.
The CDC 6600 ISA also included the instruction originally proposed by Alan Turing and implemented in the Ferranti Mark 1 computer (as "sideways add"), which was later renamed as "population count" in the Cray 1 ISA, and which was added to the x86-64 ISA by the AMD Barcelona CPUs, and later by the Intel Nehalem CPUs. It is said that this instruction was added to CDC 6600 due to a request from NSA, which then became an important customer for the CDC supercomputers, and later for the Cray supercomputers.
The usual definition for two vectors A and B is A.B = |A||B|cos ø. If A and B are binary it becomes, A and B = popcnt(A)popcnt(B)cos ø. The similarity measure is how close cos ø is to 1.
Using modern vocabulary, with a message and a set of search criteria using a word embedding on the message, calculate the similarity to each search using, cos ø = (A and B)/(popcnt A * popcnt B). This is only be vectorizable when there is an intrinsic popcnt instruction. Results that have > 0 values for cos ø mean there is some similarity.
So, to capture, (re)scan and flag all messages that contain certain keywords, design a system to generate a KEYSCORE. When eXtended to reduce false positives, build a follow on system called XKEYSCORE.
The last paragraph is pure speculation, but I think it has merit. Worth a try on your own data?
But when superscalar came around, such a RISC can execute the address update instruction concurrently with the memory access instruction.
Any superscalar CPU has limits on the number of instructions that can be fetched, decoded, renamed and dispatched and having 2 instructions instead of 1 for each load and store in an array processing loop can exceed any of those limits and slow down the execution, even when there are enough ALU/AGU execution units available.
This is why all mainstream ISAs implement either the CDC 6600 and IBM 801 solution, with auto-update of the index registers, or the DEC VAX solution with indexed addressing with scaled indices and up to 3 components (base + index + displacement). When auto-update is available, 2-component indexed addressing without index scaling is sufficient (where the 2 components may be chosen between base + index register and base + displacement).
The auto-update variant always allows loops with a minimum number of instructions, while scaled indexed addressing allows loops with a minimum number of instructions in the majority of cases, if appropriate data structures are used (i.e. SoA).