The basic thing that a programmer needs to know about parsing is that regular expressions usually aren't. The big problem with regular expressions is that they have real trouble handling nested modes. Just to give a modestly complicated example of what can happen from a simple grammar with nesting, imagine a grammar with just single-letter variables and parentheses, parsed into Python with tuples:
a(bc)def
=> ('a', ('b', 'c'), 'd', 'e', 'f')
You might be able to get a regular expression engine to deal with that for you, but it will tend to fail on something slightly more complicated like this: a(b((cde)f)(g)h)ij
...because it will usually either be too conservative and start matching: (b((cde)
or else it will sometimes be too aggressive and start matching: ((cde)f)(g)
Perl is a big exception; Perl actually has a syntax to make regular expressions properly recursive.A good practical example for you to think about is CSV parsing. In CSV, a field can sometimes be quoted so that it may contain commas, and two quoting-symbols inside that field is interpreted as a single quoting-symbol too. Thus to create a table containing a chat, we might have to write:
Reginald,I think life needs more scare-quotes.
Beartato,"That's nice, Reginald."
Reginald,"Don't you mean ""nice""?"
Beartato,Aah! Your scare-quotes scared me!
With a straightforward character-by-character parser this is pretty easy to parse; you check for the «,» character and append another record onto the row, check for the «\n» character to append this row to the table, and when you see «"» you enter a special string-parsing mode which does not exit until it sees an odd number of quote marks pass by, bordered by two non-quote marks. (The fact that you might have to "peek ahead" leads to interesting parse combinators which have to be able to succeed or fail without moving the cursor forward.)At one of my jobs I actually spent some time outside of work making CSV parsing faster by using a dangerous hack: you can split on the quote character first, to have the native string code handle the basic division of the table; then the even fields contain CSV data without quoted text (or occasionally empty strings, which must be interpreted as commands to add the quote character and the following text to the preceding field), and the odd fields contain raw data to be entered into a new record. (I call it a dangerous hack because I remember the thing working at first, but eventually I believe three bugs came out on more complicated CSV files -- it was very hard to "see that it worked right" due to the extra overhead needed to move all of this character-by-character logic into the split() calls of my programming language.