There are good historical hardware architecture textbooks: Peter Kogge’s “The Architecture of Pipelined Computers” (‘81) mentions that the UNIVAC 1 pipelined IO (from which dedicated IO units sprung out), the IBM 7094 similarly interleaved multiple-state memory access cycles between multiple memory units to increase throughput, and then the IBM STRETCH had a two-part fetch-execute without attempt to address hazards.
So pipelining came out the ad-hoc application of the patterns of interleaving and hiving off functionality to dedicated units. In general, a lot of hardware architecture ideas cropped up very early in big iron, with microcoding even co-existing with them, without all of the extra full ideas being worked out (e.g. forwarding, hazard resolution of all types, etc.)
There’s a whole fun universe of vintage hardware architecture textbooks that get through the essential concepts through historical big iron, if you’re willing to look beyond Hennessy and Patterson and Shen and Lipasti! I am thinking that the latter might have references to other historical textbooks on pipelining, though. “The Anatomy of a High-Performance Microprocessor”, a pedagogical deep-dive into the design of the AMD K6 uarch (up to detailed algorithms and HDL!), will have good references to historical textbooks too since these authors were trained before the crop of “modern” (microprocessor-era) textbooks.
1957 in a transistorized computer is far too late for the origin of pipelining.
Pipelining had already been used in computers with vacuum tubes a few years before and it had also been used already a decade earlier in computers with electromechanical relays, i.e. IBM SSEC, which had a 3-stage pipeline for the execution of its instructions (IBM SSEC had a Harvard architecture, with distinct kinds of memories for program and for data).
Before 1959, pipelining was named "overlapped execution".
The first use of the word "pipeline" was as a metaphor in the description of the IBM Stretch computer, in 1959. The word "pipeline" began to be used as a verb, with forms like "pipelining" and "pipelined" around 1965. The first paper where I have seen such a verbal use was from the US Army.
Outside computing, pipelining as a method of accelerating iterative processes had been used for centuries in the progressive assembly lines, e.g. for cars, and before that for ships or engines.
Pipelined execution and parallel execution are dual methods for accelerating an iterative process that transforms a stream of data. In the former the process is divided into subprocesses through which the stream of data passes in series, while in the latter the stream of data is divided into substreams that go in parallel through multiple instances of the process.
[1] https://en.wikipedia.org/wiki/Portsmouth_Block_Mills
The main inspiration for IBM SSEC has been Harvard Mark I, an earlier electromechanical computer built by IBM based on a design done mostly by Howard Aiken, but it would not have been impossible for some information about the Zuse computers to have reached IBM after WWII, contributing to the design of the SSEC.
With parallel do you mean superscalar execution i.e. having two ALUs or do you mean having multiple cores?
At instruction level, the execution of instructions is an iterative process, so like for any other iterative process parallelism or pipelining or both parallelism and pipelining may be used.
Most modern CPUs use both parallelism and pipelining in the execution of instructions. If you have e.g. a stream of multiply instructions, the CPU may have for example 2 multiply pipelines, where each multiply pipeline has 4 stages. The incoming multiply instruction stream is divided into 2 substreams, which are dispatched in parallel to the 2 multiply pipelines, so 2 multiply instructions are initiated in each clock cycle, which makes the example CPU superscalar. The multiply instructions are completed after 4 clock cycles, which is their latency, when they exit the pipeline. Thus, in the example CPU, 8 multiply instructions are simultaneously executed in each clock cycle, in various stages of the 2 parallel pipelines.
Whenever you have independent iterations, e.g. what looks like a "for" loop in the source program, where there are no dependencies between distinct executions of the loop body, the iterations can be executed in 4 different ways, sequentially, interleaved, pipelined or in parallel.
The last 3 ways can provide an acceleration in comparison with the sequential execution of the iteration. In the last 2 ways there is simultaneous execution of multiple iterations or multiple parts of an iteration.
For parallel execution of the iteration (like in OpenMP "parallel for" or like in NVIDIA CUDA), a thread must be created for each iteration execution and all threads are launched to be executed in parallel by multiple hardware execution units.
For pipelined execution of the iteration, the iteration body is partitioned in multiple consecutive blocks that use as input data the output data of the previous block (which may need adding additional storage variables, to separate output from input for each block), then a thread must be created for each such block that implements a part of the iteration, and then all such threads are launched to be executed in parallel by multiple hardware execution units.
These 2 ways of organizing simultaneous work, pipelined execution and parallel execution are applicable to any kind of iterative process. Two of the most important such iterative processes are the execution of a stream of instructions and the implementation of an array operation, which performs some operation on all the elements of an array. In both cases one can use a combination of pipelining and parallelism to achieve maximum speed. For these 2 cases, sometimes the terms "instruction-level parallelism and pipelining" and "data-level parallelism and pipelining" are used. The second of these 2 terms is misleading, because not data are executed in parallel or pipelined, but the iterations that process data are executed in parallel or pipelined. For any case where pipelining may be used, it is important to recognize which is the iterative process that can be implemented in this way. For unrelated tasks a.k.a. processes a.k.a. threads, only 3 ways of execution are available: sequential, interleaved and in parallel. The 4th way of execution, pipelined, is available only for iterations, where the difference between iterations and unrelated tasks is that each iteration executes the same program (i.e. the loop body, when the iteration is written with the sequential loop syntax).
You are right, but this can be make more precise: pipelining is a specific form of parallelism. After all the different stages of the pipeline are executing in parallel.
I use "parallel" with its original meaning "one besides the other", i.e. for spatial parallelism.
You use "parallel" with the meaning "simultaneous in time", because only with this meaning you can call pipelining as a form of parallelism.
You are not the only one who uses "parallel" for "simultaneous in time", but in my opinion this is a usage that must be discouraged, because it is not useful.
If you use "parallel" with your meaning, you must find new words to distinguish parallelism in space from parallelism in time. There are no such words in widespread use, so the best that you can do is to say "parallel in space" and "parallel in time", which is too cumbersome.
It is much more convenient to use "parallel" only with its original meaning, for parallelism in space, which in the case of parallel execution requires multiple equivalent execution units (unlike pipelined execution, which in most cases uses multiple non-equivalent execution units).
When "parallel" is restricted to parallelism in space, pipelining is not a form of parallelism. Both for pipelining and for parallelism there are multiple execution units that work simultaneously in time, but the stream of data passes in parallel through the parallel execution units and in series through the pipelined execution units.
With this meaning of "parallel", one can speak about "parallel execution" and "pipelined execution" without any ambiguity. It is extremely frequent to have the need to discuss about both "parallel execution" and "pipelined execution" in the same context or even in the same sentence, because these 2 techniques are normally combined in various ways.
When "parallel" is used for simultaneity in time it becomes hard to distinguish parallel in space execution from pipelined execution.
More generally, parallel in space is interesting because it is a necessary precondition for parallel in time.
One should not say that the pipeline stages are parallel, when the intended meaning is that they are separate or distinct, which is the correct precondition for their ability to work simultaneous in time.
"Parallel" says about two things that they are located side-by-side, with their front-to-back axes aligned, which is true for parallel execution units where the executions of multiple operations are initiated simultaneously in all subunits, but it is false for pipelined execution units, where the executions of multiple operations are initiated sequentially, both in the first stage and in all following stages, but the initiation of an execution is done before the completion of the previous execution, leading to executions that are overlapped in time.
The difference between parallel execution and pipelined execution is the same as between parallel connections and series connections in any kind of networks, e.g. electrical circuits or networks describing fluid flow.
Therefore it is better if the terms used in computing remain consistent with the terms used in mathematics, physics and engineering, which have already been used for centuries before the creation of the computing terminology.
Joking aside, of course, the ingenuity lies in finding a fitting solution to problem at hand, as well as knowing how to apply it.
Both ways had been used for many centuries for the manufacturing of complex things, as ways to organize the work of multiple workers.