That reminds me of how a few years back, I was able to beat the pants off all the XML parsing libs I could find using a regex in Perl. Like 7-10 times faster.
The trick is that it was a very simple XML format (Apache Lucene/SOLR), where each record was only one level deep and consisted of a set of tags of type, name and value.
The code was damn simple too. Something along these lines:
my %records
$content =~ /^.*?(?=<doc>)/gsmi;
my %r;
while ($content =~ m{<doc>(.*?)</doc>}gsmi) {
my $record = $1;
$r{$1} = $2 while $record =~ m{<(?:arr|date|str|bool|int|float) name="([^"]+)">([^<]+)</[^>]+>}gsmi;
$records{ $r{id} } = {%r};
}
I
might have been able to find a solution that was faster by configuring an XML parsing library correctly, but I suspect data marshaling would have eaten most the gains. I probably could speed it up even more by combining the record start/end with the content parsing and making a state machine, but this is damn fast already and about as conceptually simple as you can get, which has it's own appeal.
The moral is the same though, sometimes if your data is simple enough, you can easily beat industry standard solutions with minimal effor because throwing away their (for your case unneeded) flexibility can yield speed.