First memory was very limited by todays standards, say 30 to 60kb. By making lexing a separate pass, the source could be discarded before the parser started and the tokenized intermediate file written by the lexer could be read during the parsing pass. Typically, the compiler might keep the name table of all identifiers in memory between these passes.
Around the mid 70s, languages that could be compiled in a single pass were investigated. Pascal was one of these. Pascal didn’t support separate compilation either so the linking step was eliminated. Nevertheless, the early Pascal compilers still did lexing separate from parsing; I learned Pascal by studying the source for Wirth’s compiler (it had crazy inconsistent indenting).
The second reason lexing was done separately was performance. Touching every character of the input source file was a major bottleneck for compilers back then, so optimizing the lexer was perceived to be very important. By doing the lexing in a tight loop rather than being called once for ever token lots of overhead associated with these calls was eliminated.
Also, there was a questionable attraction to bottom up parsing. Knuth had shown how LR (left to right) shift-reduce parsers could parse a very large family of grammars efficiently around 1965. Everyone was enamored with them. I even wrote a set of FORTRAN programs that would construct the SLR tables suitable for a subset of LR parsable grammars. When lex and yacc applications came along, everyone thought that every compiler should be built this way (to be fair, Wirth and Per Brinch Hansen were both designing languages that could easily be parsed by recursive descent because their grammars were LL(1) a smaller family of grammars than LR(n)). I’ve never seen anyone try to use only a LR parser at the lexical level combined with the normal grammar parsing.
I went back to University for another graduate degree in 1984, and I was surprised to hear the professor that taught the compiler class say that everyone should be using lex and yacc for any compiler development. By then, in the real world, people had discovered that much more meaningful error messages were possible with LL or recursive descent (i.e. top down) parsing.
Now, performance and memory considerations are different. It’s practical for compilers to read an entire source file in a single read (or memory map the entire source file), saving all the round trips to the OS for reading input. Top down parsing provides better error messages and it fits well with parsing all the way down to the lexiems.