Regexes Parse XML Just Fine, Actually
evincarofautumn.blogspot.com
evincarofautumn.blogspot.com
What are you going to do with the output of this script? If you want to do anything useful with XML, you usually build a DOM out of it during the parse.
What do you mean by this? "I would call this a good ... practical solution to a problem that’s theoretically unsolvable." I'd say that parsing XML is definitely a solvable (and solved) problem.
I am getting a little tired of these types of responses to "you can't parse X with regular expressions." Someone always feels the need to come up with a tremendously complex solution that mostly works. Usually using backref matchings, (making the regular expression not so ... regular). Meanwhile, the author never seems to talk about the parse time. Which, will be exponential. See http://swtch.com/~rsc/regexp/ for all the glory details.
Specific points:
> I’m talking about parsing XML, not checking whether some input actually is XML. Correctness is a Boolean, after all: invalid XML is not XML.
However, from a particularly pedantic point of view: To parse is to check whether the input is in your language. If your parser excepts input that is not XML it is not an XML parser. It is a parser which accepts some language which largely overlaps with XML, but is not XML.
I know it's an academic point, but it is an important one. When you have properly parsed your input, you should be sure that the input is what you were expecting. A parser that excepts a different language than the intended one can be highly misleading.
So I am left to wonder what is the point of it all. You never say in the article.
tl;dr
Read the dragon book. Learn what a Language actually is. Study some Chomsky. Write a recursive descent parser. Lay in the green green grass. Think about Kurt Gödel. Write a parsing framework for LALR grammars. Think. Then, I believe, the desire for using regular expressions for in-appropriate purposes will have left you. Your other tools are just so much cooler.
Example: I've tried two parts of the xml syntax, <?xml version="1.0"?> and namespaces <foo:bar />, and it didn't seem to parse either. (At least if I understood the output right, it doesn't say "parsed" or "not parsed").
So it's the typical "author's idea of xml" parser, not xml parser.
The point is that you probably don’t want to do this with regular expressions, even though, contrary to popular opinion, you pretty much can. The article even calls the solution “slow and overspecified, but still kinda neat”.
http://stackoverflow.com/questions/1732348/regex-match-open-...
For example: write a regular expression that will return the contents of the first <div> that is found in the following XML:
For example: <div> whatever <div> more stoff </div> foo </div>
or
<div> whatever foo </div>
You can't because XML is not a regular language and regexes only match regular languages.
Also, the article does say that you can parse well-formed XML with a single regex, you just can’t validate it that way.
You missed the part about the first job of a parser being "check for correct syntax":
http://en.wikipedia.org/wiki/Parser
Parser
In computing, a parser is one of the components in an interpreter or compiler, which checks for correct syntax and builds a data structure (often some kind of parse tree, abstract syntax tree or other hierarchical structure) implicit in the input tokens. The parser often uses a separate lexical analyser to create tokens from the sequence of input characters. Parsers may be programmed by hand or may be (semi-)automatically generated (in some programming languages) by a tool.
Everyone who thinks that http://www.w3.org/TR/xml/ can be parsed with a few lines of perl is just mislead. Example: Does the parser handle XML namespaces?
At best it's a trickier way of searching xml files.
even validated HTML isn't XML, let alone tag soup.