The reason s-expressions are better than text is that they are more expressive. They allow arbitrary levels of nesting whereas text doesn't. Text naturally develops a record structure with only two levels of hierarchy: records separated by newlines, containing fields separated by some other delimiter (usually spaces, tabs or commas, sometimes pipes, rarely anything else). Going any deeper than that is not "native" for text. That's why, say, parsing HTML is "hard" for the unix mindset. But a Lisper naturally sees HTML (and XML and JSON and, well, just about everything) as just an S-expression with a more cumbersome syntax.