Two reasons, one questionably relevant and one difficult to formalize.
The questionably relevant reason is that all well-known classes of declarative syntax specifications that are easy to parse don’t really compose very well, in any direction; the exception being PEGs, which do compose, but in nonintuitive ways, and are also kind of difficult to parse. So if you want a generated parser, chaining a regular-language-based lexer and an LL- or LR-based parser lets you tackle a much wider class of languages than the second step alone.
Of course, hardly any serious language implementations use generated parsers these days, and hand-rolling a scannerless parser is not that much more difficult as far as the mechanics go.
But then, the difficult-to-formalize reason is that people do actually think of languages that way, and separating phases gets you more localized and more understandable errors. This might be an artifact of how language specs are written, but I don’t think so: good frontends will occasionally split the process into even more phases, first parsing a looser syntax than specified and then tightening it up with semantic checks. For example, ISO C prohibits (a + b = c) at the grammar level, by restricting which expression productions are allowed to the left of an assignment, but a good compiler will parse this successfully and only afterwards tell you that (a + b) is not a thing you can assign to. In this connection, the "skeleton syntax tree" idea as used in Dylan[1] seems promising, but I don’t think I’ve seen it developed further.
[1] Bachrach, Playford, "D-expressions: Lisp power, Dylan style", https://people.csail.mit.edu/jrb/Projects/dexprs.pdf