Getting the variable names used for the caller is of course much harder, and meaningless for a call like f(1+g(b)). py.test, which made a more serious attempt at handling the f() case, gave up on that approach, preferring to rewrite the original code behind the scenes rather than infer it backwards.
So you're a troll. Thank you for letting me know. By trolling the uninformed you seemingly have no meaningful point other than to point your finger and laugh at the views of others. This action runs counter to your stated goal of wanting scientists to be better programmers. What good do you think it will do to the world to troll people using your novice Python skills?
You do realize that when someone searches for "Matt Oates", "bioinformatics" and "Python" they are likely to find this page in the future, and learn about your love of trolling, right?
I do not want to live in a troll world, just like I don't want to live in a world where people use false and undercutting statements disguised as joviality. Instead, I teach scientists how to be better at the programming they need to do good science.
Your last paragraph tells me that you have very little experience in writing parsers. In a record oriented format, and with hand-written parsers, resynchronization to the record boundary is easy. In Perl's it's almost trivial, because nearly all record-oriented formats have a boundary which can be handled with $/ ; FASTA is an exception but still easy to handle. On the other hand, my point is that it's much harder to resynchronize with a context-free grammar.
Here's an example of a failing case with my strict parser. HAS2_CHICK was the only record in SWISS-PROT 39 with a DT line of:
DT 30-MAY-2000 (REL. 39, Created)
This failed because the spec said that the "REL." was only supposed to be a "Rel." My parser failed. Every other parser basically said (by lack of full validation) that they weren't interested in that field, and ignored it.
You can certainly condemn the entire SWISS-PROT 39 release as "garbage" for having a single record which didn't match the spec. Effectively everyone who wants the data will disagree with you. You can also train your parser to ignore those details, and retrain it for every release, though you'll find it increasingly difficult to accept the format as it changes across 5 different database releases.
Here's another example. Some of the GENBANK records included a qualifier of "/primer", which wasn't in the specification. My parsers failed, because it validated the qualifier names. And you know what? No one really cares about that failure, so I had to tweak my parser to allow the non-standard name. The same with EMBL: the ID line is supposed to have the molecule type (one of "DNA|RNA|circular DNA|circular RNA"). In actuality? Record S40706 had a type of 'XXX'... and so again, my validating parser stopped, when in reality almost no one cares.
As these are further examples of simple differences causing a failure within a record, your bringing up "completely arbitrary and unknown problems with a file" suggests to me that you have little real desire to engage in my point, and rather prefer to troll me by diverting the topic. In case I am wrong, please assume that Perl's $/ or a similarly simple method is sufficient to determine the record boundary and resynchronize.
The underlying issue is that a database format spec changes after almost every release, so a validating parser ends up having to handle a family of formats, not one. Which is much easier for a line-oriented parser to do, since those formats are written with line-oriented hand-written parsers in mind, and usually updated in such a way as to not break those parsers.
When you find out that your nice, elegant, validating context-free parser breaks after nearly every release, due to details that are irrelevant to what you care about, then you might conclude that the rest of the world is a bunch of ignorant fools who can't program; or reject the idea of a validating parser for this sort of data and switch to a line-oriented one; or provide a recovery mechanism to determine what to do with the records which couldn't be parsed, and continue if requested. I do the latter now.
(It's also difficult to use a CFG-based parser that doesn't know about record boundaries to parse files that don't fit into memory even though each record does. I don't know if that's still an issue in bioinformatics these days, but it certainly is in cheminformatics. In any case, a mix of simple record extraction and a CFG-based grammar suffice, excepting for records containing an entire genome, which is outside the scope of this sort of task, and for PIR, which had header and footer components to the file.)
Regarding performance, don't trust your beliefs or mine. Measure it. Empiricism trumps faith. Try benchmarking a hand-written, line-oriented FASTA parser against your grammar. I wrote one just now in Python. The regex which used "[^>\n].*\n" was about 2% faster than checking the [specific sequence letters]. A 2% here and a 2% there and those validations quickly incur noticeable overhead. By comparison, my line-oriented parser took about 1/2 the time of the regex-based parser.
Of course, Perl has a different runtime environment and will have different conclusions.
Then again, there's already a history in Perl of variable stringency levels for improved performance. Swissknife (doi: 10.1093/bioinformatics/15.9.771) was written as a lazy parser because Perl5 wasn't fast enough for full validation and parsing of each record. Partial evaluation gave a 4x performance improvement when full validation wasn't needed.