The Parser that Cracked the MediaWiki Code
dirkriehle.com
dirkriehle.com
http://www.mediawiki.org/wiki/Alternative_parsers#Known_impl...
Most of these are special purpose hacks. Kiwi and Sweble are the most serious projects I'm aware of, that have tried to generate a full parse.
However, few of these projects are useful for upgrading Wikipedia itself. Even the general parsers like Sweble are effectively special-purpose, since we have a lot of PHP that hooks into the parser and warps its behaviour in "interesting" ways. The average parser geek usually wants to write to a cleaner spec in, well, any language other than PHP. ;)
Currently the Wikimedia Foundation is just starting a MediaWiki.next project. Parsing is just one of the things we are going to change in major ways -- fixing this will make it much easier to do WYSIWYG editing or to publish content in ways that aren't just HTML pages.
(Obviously we will be looking at Sweble carefully.)
If this sounds like a fun project to you, please get in touch! Or check out the "Future" portal on MediaWiki.org.
Did you ever turn up anything regarding this? http://news.ycombinator.com/item?id=2216249
btw, neat js parser, i'll have to check it out. :)
As for your original question, I don't think there is a forum that tries to unite the left-brained and right-brained wikipedians. There is a bit of a divide. I'll send an email right now to someone who might know better.
We don't have contests per se to try to steer the community, other than I guess GSoC, or reaching out to developers that we think are already doing good things.
For one, sweble is a Java parser, and I'm not sure this makes it a good drop-in replacement for the current MediaWiki PHP code. The DBPedia Project also has what looks like a decent AST-based Java parser [1]. I would be interested in a comparison between sweble and DBPedia's WikiParser.
I stumbled across a very nice MediaWiki scanner and parser in C a while ago [2]. It uses ragel [3] for the scanner; the parser is not a completely generic AST builder, but is rather specific to the problem of converting MediaWiki markup to some other wiki markup. It does do quite a bit of the parser work already though.
Presumably a PHP extension around a C or C++ scanner/parser could someday replace the current MediaWiki parsing code.
[1] http://wiki.dbpedia.org/DeveloperDocumentation/WikiParser?v=...
Of course, a PHP parser would still have to be maintained in parallel as not everyone would be able to do the Java option.
No doubt the incremental complexity for Wikipedia would be small in relative terms. I assume that argument would support a variety of proposals.
A solid scanner and parser in C/C++ would benefit a broader audience though. All the major scripting languages can be extended in C/C++. In fact, the ragel-based parser I mentioned earlier [1] was built to be used from within Ruby code.
cache version: http://webcache.googleusercontent.com/search?q=cache:8xjwEj-...
http://dirkriehle.com.nyud.net/2011/05/01/the-parser-that-cr...
"It worked for me." ;-)
First, I believe this reveals the complexity of the parser, which implies a complex syntax, which implies a complex user interface as felt by end users. A more complex the user interface may make it harder it is to attract new editors, although it's unclear (to me) if that is a fact.
Second, having an AST representation is awesome. It makes it possible to even think about building a path towards WYSIWYG or some other form of rich text editing. It was not really possible to build a WYSIWYG editor around the wiki syntax.
If you have an AST, you can also store the page as the AST since you can regenerate the wiki syntax from the AST for people who need text-based editors.
There may be friction against gaining new editors from the user interface of the MediaWiki software, but I think the greatest barrier to participation by new editors is the hostile, drama-filled environment on many controversial topics on Wikipedia. My evidence for that is the decline in "unsustainable fashion"
http://strategy.wikimedia.org/wiki/Story_of_Wikimedia_Editor...
in the number of Wikipedian administrators, who presumably for the most part are people who know how to use Wikimedia software. Too many of best contributors (people who look up facts in reliable sources and edit articles for better readability) on Wikipedia feel attacked and that their time is wasted. I know a lot of dedicated hobbyists who quietly work on their hobby-related subjects putting together great articles, but on any subject that is controversial, and for which looking up reliable sources takes some effort, Wikipedia is becoming a war zone and is not improving in quality.
http://strategy.wikimedia.org/wiki/Strategic_Plan/Movement_P...
http://strategy.wikimedia.org/wiki/Strategic_Plan/Movement_P...
That's the case to some extent, but the opposite is also the case to some extent. Some of the difficulty of parsing is because "ease of human use" has been a much higher priority than "ease of parsing" when discussing syntax, which leads to some constructs that aren't easy to parse with typical CFG-type parsing approaches. It's also designed to be very lenient to ordering and common errors, much like a modern non-strict HTML parser, which makes hand-writing the syntax more friendly and forgiving, but with a tradeoff that the parser has to be more complex, because it doesn't have the luxury of just returning a parse error.
I suppose this is one of the knobs that must be tuned to balance between reproducible I/O and turning away meaningful contributions from the community.
If it was anything else, I'd say "who cares," but this is "the world's knowledge" -- we absolutely should care about the format it's stored in. I'm glad to see people tackling this problem.
Another example of the imperfect but working implementation winning.
For parsing MediaWiki in Python, check out mwlib [1], which was part of cooperation between Wikimedia Foundation and PediaPress. It's neither very complete nor very fast, but you might be able to hack up some LaTeX conversion with it.
Personally I'd be happy to see any markup language becoming the default, regardless which one it is. Having a proper grammar would be a bonus.
I don't agree with this. All lightweight humane markup languages are not born equal, some are better others, and Mediawiki's is not in the list of the best ones. Now there seem to be a trend towards Markdown, but it should be improved and then, migrating Wikipedia to this Markdown2 could be a real good thing.
I actually wrote a subset-of-MediaWiki parser in C#, which stayed at the "subset" stage because of the ridiculous complexity and corner-cases that crop up even very early.