What Lua scripting means for Wikimedia and open source
blog.wikimedia.org
blog.wikimedia.org
As far as I know, something like this has not been tried by anyone else at this scale before.
Quoting from the original post: "Anyone can write a chunk of code to be included in an article that will be seen by millions of people, often without much review. We are taking our 'anyone can edit' maxim one big step forward."
The debug lib, allowing to take control of the whole program, is disabled in most sandboxed environments including wikipedia's.
However, bytecode instruction and line changes both are poor metrics to measure time, cpu or any other kind of resource usage. Measuring process time is relatively easy and much more accurate.
I don't think that's quite true. MediaWiki templates are now more powerful, but I'm not sure even that counts as interactive. This makes editing easier, but the reader isn't going to interact with a Wikipedia page any more than they did last week.
If one thing can be said about Lua, it is that embedding Lua is a solved problem, and not something dangerous to be worried about.
1. Every page is now a "real" program. Will separation of code and content result from this? The article hints that structured data is on the way, but I have a hard time imagining the average contributor will restrain themselves given this awesome new power. Instead I foresee bigger and badder messes.
2. Does this get us any closer to a parseable format for Wikipedia content? This is the world's knowledge. The fact that it's currently trapped in an ad hoc, PHP-inspired template "language" (there is no formal grammar for it, it's whatever some PHP code accepts as input) should be very concerning.
3. Why Lua? Javascript seemed like the obvious choice. Rendering templates server-side at Wikipedia's scale sounds incredibly challenging; it boggles the mind really. Surely most page views are from js-enabled browsers -- why not offload all that rendering to the client?
https://www.mediawiki.org/wiki/User:Sumanah/Lua_vs_Javascrip...
Wikipedia pages are massively cached: rendered pages are cached in Squid, and I believe parsed page contents are stored in memcached. Templates are only re-rendered when pages are edited, if all goes well.
Transforming wikitext into HTML is surprisingly IO-bound, since most Wikipedia pages transclude many other pages for templating and logic. It's very slow now, and you need to be close to the database for this to be remotely practical.
Lua gives the community a way to replace certain horrible templates which perverted transclusion and macro-expansion to do logic. But most Wikipedia pages will still transclude a lot of other pages.
Also, web browsers are not the only consumer of Wikipedia pages; you can also download PDFs and so on.
Sounds like a case for in memory caching. Or is it that any random Wikipedia page requires logic from a random subset of a large pool of other pages? Somehow I doubt that, with Zipf's law coming to my mind.
Also remember that an HTML5 front-end would be able to cache a substantial amount of this (a few megs at least) and Lua can be quite effectively compiled into JS.
For a sense of how hard this would be, try using the Special:Export page on Wikipedia.
http://en.wikipedia.org/wiki/Special:Export
If you download transcluded templates, the article for Barack Obama is 781K.
My experience of Special:Export is that it has some flaws that cause it to miss some things it needs to export, so the real total may be much larger.
And that's just the data - one would also have to download a lot of related code, which might balloon that up to a megabyte or more.
Wikitext is particularly ornery (because it's just based on grinding regular expressions against each line, it is not easy to describe with regular grammars) so you'd have to download a very large parser, with various plugins as each MediaWiki install uses them to warp how Wikitext is processed. This is assuming some optimistic scenario where MediaWiki's rendering, and all related plugins, are entirely ported to JavaScript compatible with all desired browsers.
I'm not denying that, if you wanted to create a new Wikipedia from scratch today, based on JavaScript, you could probably move a lot of rendering to the browser. You would choose more browser-friendly formats, like JSON or XML, rather than making up some random text-based format, just because it was easy to type into a textarea. You would make transformational operations work in JavaScript, or be exportable to JavaScript. You could definitely get it to the point where it would be practical for quick previews while editing.
For the Wikipedia that we have today, it's really hard.
?! I can play Quake in Javascript.
I have been running Node.js and Lua at fairly large scale on EC2 recently.
I like the V8/Lua/Luajit platforms just fine. Each has some nice strengths.
With Lua, I see significantly less memory use, and moderately increased performance.
As long as RAM is the expensive bottleneck for hosted servers, it looks like I will be able to get better performance (in operations per dollar) by using a smaller runtime.
Why not client-side rendering?
Accessibility? Cross-platform uniformity? (better answer from 'capnrefsmmat: render once, then cache massively)
http://thread.gmane.org/gmane.science.linguistics.wikipedia....
Wikimedia needs to move beyond 1990s internet and allow interactive articles. Articles on 3D that have interactive 3D. Articles on Physics that have interactive simulations. If you asked me in 1980 what an encyclopedia would look like in 2013 I would not have guessed static text + static pictures.
I know it's not an "either or" thing. They can have both. But I think once they get around to adding JavaScript they'll regret having added Lua now and having 2 code bases with 2 sets of libraries and code to synchronize the 2 when they need to communicate with each other.
What I'm sad about is that JavaScript is so locked-in on the browser that I can't use Lua there as well.
As for fast, I would love to see more comparisons on real-world code, there are too few on the web.
It is easier to embed Lua than JavaScript, however, that's say a week of work, then you're done, so it isn't a reason for something like Wikipedia to prefer Lua.
With all that said, Lua and JavaScript are both good languages, adding either one to Wikipedia is going to be a huge improvement.
edit: looks like my phrasing has offended mikemike, who I have the utmost respect for, so I removed some stuff he disliked.
Only to programmers raised in a C-based world. Normal people, mathematicians, and Lua start counting at 1.
I remember when I was new to programming I was horribly annoyed by the fact that C started counting at 0 - that was so unnatural! (the BASIC I had used up until that point used 1-based indexing too)
0-based indexing is good for low-level code, for memory and screen addresses. For high level concepts starting at 1 is the natural choice.
However, I admit after many years of C/C++ I was quite irritated by everything starting at 1 by default. My brain adapted after a few days, though.
I think the Lua devs made the right choice with the 1-based indexing. Just like 0-based indexing was the right choice for C. As I said both ways are natural but in different scenarios: 0-based for low-level machine code, 1-based for high level human concepts.
I mostly agree, except for mathematicians ;) , where it depends on the context. For example aleph numbers start at 0,
http://en.wikipedia.org/wiki/Aleph_number
and even the natural numbers can start at 0 or 1,
Large standard libraries would mean huge potential for security holes in this case; they've locked features down to only support things that they could make secure.
And starting from 1 you can get used to. LuaJIT doesn't even REQUIRE you start from 1; it'll optimize arrays you start from zero. And even vanilla Lua will accept 0 as a table index; it just won't treat it as an array (for ipairs() or length).
But if you start asking people for complaints about JavaScript, you will find people complaining about operators acting wonky, auto-type-conversion problems from hell, and even odder things like [1].
Lua is just a cleaner design. I am stuck using JavaScript when I want to work on the web, but knowing what an elegant language CAN be makes me enjoy it a lot less. And don't get me started on how ugly NodeJS is because JavaScript doesn't have coroutines...
Watch the YouTube video of the Wikipedia presentation linked elsewhere in this discussion.
From the video, they picked Lua also because they could have multiple concurrent Lua states (no globals), easy to override the default memory allocators for security, Lua VM instances are very cheap and fast to create and launch, and the code base was small enough that their security guy could actually audit all the code (and they even submitted a patch).
All these properties are hard to find with Javascript, but with Javascript, you have to go further and first decide which Javascript implementation to use. The ECMA spec says nothing about C interfaces or embedding so every Javascript engine (e.g. V8, JSCore, SpiderMonkey) is completely different.
They also said they are using the standard Lua interpreter and not LuaJIT because stock Lua happened to already be really fast and fast enough for their purposes right now.
I didn't say otherwise, I know LuaJIT is used. What I did say was I have not seen speed comparisons on such real-world code - but I would love to see a link if you have one.
> and the code base was small enough that their security guy could actually audit all the code
A small codebase is indeed good, but JavaScript VMs have very large teams of people working on hardening them, both reading and reviewing the code and applying methods like fuzzing to look for security problems. The fact is that JS VMs are likely the more hardened of anything out there, simply because they are used in web browsers.
> with Javascript, you have to go further and first decide which Javascript implementation to use.
Lua has at least two implementations as well, mainstream Lua and LuaJIT.
> They also said they are using the standard Lua interpreter and not LuaJIT because stock Lua happened to already be really fast and fast enough for their purposes right now.
That's cool.
The difference is that the Lua C API is standardized meaning that C code can work with either just fine. The only big incompatibility problem is that LuaJIT provides an extra, non-portable, FFI library in addition to the standard one.
Deleted comment
I can find that link to the v8 blogpost, if you want?
edit: While I didn't mean them the way you interpreted them, I edited out the parts of my previous comment that you took offense to.
LUAJIT_ENABLE_CHECKHOOK
edit: ditto.(When doing things like "given an interval of indices a-to-b, partition it into n subintervals so that every element of the original interval is in precisely one of the sub-intervals, and the lengths of the intervals are either the same or as close to each other as possible", I find the "count from one, bounds are inclusive" model horrible. How do you express an interval from a with the length of zero? It's just <a,a) with half-open intervals. What, <a, a-1>? So many special cases with the <a, b> model...)
Static text + pictures is exactly the way I like my encyclopedia to be. It is simple, lightweight, works everywhere, printable, usable on e-ink displays, and does the job just fine. You can always link to interactive toys, but please keep the main articles this way.
They are more resolving server performance issues. The client processing is not to be affected.
http://tests.web2py.com/physics2d/default/code/10
That points out that arguably Wikipedia will eventually have to allow JavaScript in pages it serves if it wants to progress.
When that happens there will now be 2 code bases. The Lua code that generates pages and the JavaScript code that does interactive examples for articles.
So the question is, why have 2 language which requires twice as many libraries, twice as much knowledge, twice as much expertise, and various glue to get them to interact with each other, vs having 1 language used for both?
One perspective: Because choosing a language because it is conveniently in place should not be the primary reason. Many people are comfortable with JavaScript, but there are many server side code bases in other languages because they are preferred for various reasons.
There is a large amount of PHP code deployed, but there does not seem to much drive or concern that it does not run in browsers. Both JS and PHP earn a large amount of negativity because of their shortcomings (which we don't need to go into here), and they are obviously not the tool for every job.
I would simply make the point that choosing a language to avoid having to learn more than one is not an approach for improvement. Otherwise as I said in another post, we would all be using IE with ActiveX controls.
If the code is in a different language then both of those are often false. The 2 programmers don't speak the same language and even if they both happen to understand both languages they can't share code.
JavaScript did not birth with a fully formed userbase.
The whole idea behind this effort is to facilitate contributions to Wikimedia from a diverse set of (presumably busy) contributors, correct? If so, what's the point in making them learn a new language in order to contribute, when an existing well-known language is a perfectly reasonable fit?
Sounds like someone had a Lua hammer, so everything looked like a Lua nail.
Same could be said for most of the JS proponents in the thread.
https://en.wikipedia.org/wiki/User:Dragons_flight/Lua_perfor...
Now, if we could only go back in time and replace Vimscript with Lua or Lisp or anything...
That sounds like the life story of PHP! And many other similar solutions.
"People who have programming skills should enjoy better productivity and lower frustration while editing Wikipedia. Perhaps most exciting is how many people will have a gentle and practical introduction to programming because of this."
A lot of people on HN get down on Wikipedia because of deletionists etc. and support a forked version but the fact is Wikipedia is the most salient expression of the spirit of the internet. Google is great but is a business with all the downsides that entails, project abandonment etc. Maybe I am gushing but Wikipedia is without peer.