What's New in Lua 5.4
lwn.net
lwn.net
It's 5.1 plus some of 5.2.
[0] http://www.opengroup.org/comsource/techref2/CHP06GDC.HTM
Very much an XKCD 927 story.
[0] https://en.wikipedia.org/wiki/Raft_(computer_science)
[1] https://en.wikipedia.org/wiki/Log-structured_file_system
He was a smart and productive person that designed an ugly programming language.
Lua https://p.thorsen.pm/5ecb8ce026ab
Tcl https://p.thorsen.pm/ee8a792acd52
Results: Lua 5.2: 0m0,027s
LuaJIT 2.1: 0m0,016s
Tcl 8.6: 0m0,630s
I'd pick Lua over Tcl any day.I got an order of magnitude improvement on Tcl just by bracing the expressions:
Bad: expr $i +$j + 2
Good: expr {$i + $j + 2}
... it saves interpretation and evaluation to occurring once, hoisted in the [expr] engine; it’s idiomatic. The lua version still appeared faster, though I’ve only spent 2 minutes looking at this.
Again, strings should play no issue here; values must be able to generate a string-representation, but it’s computed lazily, and in this case, the $i, $j, etc in Tcl are working with native integers.
I’m happy we have things like Tcl and Lua; either one of them makes development enjoyable. I appreciate you taking time to write your code samples.
Keep in mind that while, yes, Tcl lists have a more efficient internal representation when Tcl knows that it's a list, a string that happens to be a valid representation of a list is also a list, just by virtue of existing - but it won't get a special list representation until you try to use it like one! It's just a string like "foo bar baz". Or, say, "objref#1 objref#2 objref#3".
And the semantics of the language is EIAS through and through - all those internal representations are optimizations that are not supposed to change the observed behavior, only performance. A GC cannot treat lists without internal representation differently from those with it without breaking this. So it has to assume that any string is potentially a list that contains object references.
So a proper Tcl GC would have to be incredibly pessimistic about what it considers to be roots of the object graph - basically, any substring that looks like an object reference, has to be considered as one. Now imagine just how much string scanning such a GC would have to do in a non-trivial app for a single sweep!
It sounds like you’re familiar with Tcl internals, so I presume you know that (if we stick w the list as a prototypical “complex” structure) that the management is reference counts, freed when count==0, and may be >1 if this is a shared item (I’m going to refrain from calling it an object, though internally they’re a Tcl_Obj type, which has nothing to do with OO-programming). In a list, the list is an item, and ea. element is an item, with (of course) incr/decr references at when appropriate, as determined by the language rules itself.
Wrt strings vs lists, a string could be considered a single “blob”, but coercion to a proper list (“shimmering” in Tcl parlance) does the tokenization with requisite ref. counting as part of the conversion.
...I think the above may be entertainment more for anybody else who’s following along, but if that illuminating for you too, that’s a bonus.
Perhaps you could show me by example:
What’s a piece of (eg) Ruby and what you think a Tcl workalike would appear as, and where that falls down.
(I'm oversimplifying and ignoring the command part of it, because it doesn't really matter for any of this.)
For the sake of readability, let's assume that those object IDs / namespace names are just numbers, e.g. ::42 (in practice, it's something like ::oo::Obj42). Since that number was generated by TclOO machinery under the hood, said machinery can properly tag it as a reference at that point, and return such tagged value from [... new].
Now, let's say that we have some class, and some code that constructs an instance of that class, and then builds a list that references that instance from one of the elements. One way to do it would be start with an empty list, and build it element by element:
set foo [Foo new]
set lst {}
lappend lst blah
lappend lst $foo
lappend lst blah
unset foo
This produces the list {blah ::42 blah}. Because we built it element-by-element, our Tcl-with-GC tags it as a list, and uses an optimized representation for it, in which ::42 is a separate element that's known to be an object reference. Even after we remove the direct reference by unsetting foo, GC can use that metadata to trace a path to ::42 via lst, and know to keep the object alive. If we want to recover foo later, we can safely do: set foo [lindex lst 1]
So far, so good. But... that wasn't idiomatic Tcl. We're much more likely to build the list thus: set lst "blah $foo blah"
unset foo
It's still the same list {blah ::42 blah}. But, since we haven't performed any list operations on it, Tcl had no reason to tag it as one, and to switch it to the optimized representation - it's really just a string at this point. And because the elements haven't been separated out, we've also lost the "this is an object reference" tag on ::42 - there's no separate element as yet, so there's nothing to tag!So then, how would GC know to look at lst, treat it as a list, and find the object reference ::42 inside? Note that if it doesn't do that, then the object will possibly be garbage collected by the time we get to:
set foo [lindex lst 1]
Tcl finally realizes that lst is a list because of our use of lindex here, and converts it to the efficient representation that separates the elements... but at this point, how would it know that the element ::42 originated as an object reference, rather than a random string that looks like one? And even if it just pessimistically assumes that it's a reference, that doesn't really help - the object is already collected, and the reference is invalid anyway. So, we just made shimmering observable, and thus broke EIAS semantics. Our Tcl-with-GC is no longer Tcl.So, to remain true to EIAS, our GC can't rely on the presence or absence of efficient list representation for a value to decide whether to scan it for references, or not. It has to proactively treat every variable in the program as a potential list with potential object references in it - because, well, it could be; and if it is, then it's also a GC root that must be traced!
And yet this is a trivial case. We actually have to consider nested lists, and dicts, and unevaluated Tcl code stored in a string, and all possible permutations thereof...
This is how the garbage collector in Jim Tcl works. It scans the string representation of all objects [1] for the reference syntax [2]. In practice, it does not cause performance problems. Since Jim Tcl is intended as a smaller, more easily embedded counterpart to Tcl 8, nobody runs it with a large heap.
How slow is it? I have just tried a naive benchmark on a first-gen Raspberry Pi Model B. With a live list of one millions strings and 50 thousand references garbage collection took around 750 ms. On a Core 2 Duo laptop running Linux it took around 120 ms.
. for {set i 0} {$i < 1000000} {incr i} {
set v$i "<reference.<--NOT padding padding $i"
}
. for {set i 0} {$i < 50000} {incr i} {
set r$i [ref $i test]
}
. time collect 5 ;# Average of five runs
765443 microseconds per iteration
The tracing garbage collector only affects references. Normal strings are reference-counted like in Tcl 8. References are the low-level basis of Jim's OO system, which is implemented in pure Tcl.[1] "Objects" in the sense of Jim_Obj (Tcl_Obj). https://github.com/msteveb/jimtcl/blob/da293a1eef2ddd709f10b...
[2] http://jim.tcl.tk/fossil/doc/trunk/Tcl_shipped.html#_garbage...
I should have just said "all values". The name Tcl_Obj/Jim_Obj is a legacy artifact that leads to lots of confusion. Perhaps Tcl 9 will rename Tcl_Obj to Tcl_Value.
Aesthetics matter.
Unfortunately, LuaJIT is tightly-coupled to the Lua language and its code is very complex. I'd be surprised if anyone other than Mike Pall was able to retarget it to support a different language.
Also, is JIT compilation really necessary for a scripting language? Projects like the PyPy JIT for Python have seen little adoption, and many platforms (e.g. iOS and game consoles) don't support JIT at all.
Wow, the understatement of the year.
Not that difficult. See e.g. https://github.com/rochus-keller/Oberon#lua-source-and-bytec..., https://github.com/rochus-keller/Smalltalk#a-smalltalk-80-in..., or here a whole list of other language frontends: https://github.com/hengestone/lua-languages
The projects you link to all either compile to, or are written in, Lua source code or LuaJIT bytecode. That's less efficient than retargeting the LuaJIT VM. For example, LuaJIT includes bytecode instructions specifically tailored to Lua tables. It looks like these projects build, say, their array data structures on top of those existing instructions, rather than adding array-specific support to LuaJIT.
That's like saying LLVM is inefficient because all the frontends compile to LLVM IR.
E.g. the Oberon compiler directly generates efficient LuaJIT bytecode also using FFI and even reuses the VM features for it's source-level debugger. Then benchmark written in Oberon has nearly the same performance like compiled to native when it runs on LuaJIT
1. Decouple the extension mechanism. Use cffi
2. Modularize the batteries
Each alternative Python has to re-implement large swaths of the stdlib. The code is the spec in Python.
Stackless got the PyPy treatment before PyPy even existed. Besides PyPy, we now have graalpython, rustpython and ...
https://github.com/beeware/ouroboros/issues/21 points to a python stdlib I had never heard of, https://github.com/pfalcon/pycopy-lib
The point is, the core CPython devs need stop being so selfish with their toys. The greater Python ecosystem is diverse in spite of their actions not because of it.
I was under the impression that LuaJIT has support for at least one game console because it's used in games. Is that not the case?
"Due to restrictions on consoles, the JIT compiler is disabled and only the fast interpreter is built." [0]
I moved from Matlab to Python and the biggest improvement was zero-based indexes.
Apparently this little detail is never taken into account.
You don't if you're writing low level algorithms. It's difficult for me to articulate the exact problem but I liked 1-based indexing until I needed to implement some algorithms that operated on multi-dimensional arrays element wise. Things just don't compose the right way.
It's actually the primary reason I don't make more use of Julia. I like the syntax and overall design more than Python but it just isn't worth it to me.
(If you're wondering what I'm doing in an Lua comments section, it's because I really like the ideas behind some of the Lua based languages such as Terra.)
One or zero indexing is a question about domain and communication. In signal processing, it is quite normal to need a negative index (autocorrelations f.x.). Ada (and I believe Pascal and Julia) has supported for specifying the indexing type, which can then be selected to match the problem domain.
Maybe you should try it before passing judgement.
In LUA you can use the metatable functionality to implement 0 based indexes.
In Julia you can overload the array access operation too. There a packages which allow zero based indexing. Or even StarWars based indexing including the machete viewing order. https://github.com/giordano/StarWarsArrays.jl Julia's array interface has has functionality what the first/last index is, so you can use it in other peoples code. If they haer coded 1 as first index just make a pull request to fix it.
I tried to do that a while back and kept coming across corner cases that didn’t work. Is there a complete implementation of zero-based arrays for Lua available anywhere?
I've had an ok using unicode libraries but I haven't had to do much with them so far.
Imagine if Kubernetes config files are actually Lua files.
Thanks for this mention - I found:
CloudFlare repos on GitHub using Lua: https://github.com/cloudflare?q=lua
Blog posts with "Lua" tag: https://blog.cloudflare.com/tag/lua/
Pushing Nginx to its limit with Lua - https://blog.cloudflare.com/pushing-nginx-to-its-limit-with-...
You just have to practice self-discipline and not let Turing creep take over.
- global variables by defauilt
- weird choice of certain operators (~= for not equal)
- arrays start at 1
(fixed yet?)May work better for mathematicians, non-SW engineers etc, but I wouldn't know.
It may not be a language that website jockeys use, but "I don't use it therefore probably no one does" is almost never correct. MATLAB is huge in state machine and dynamic systems programming as well as being the original commercial powerhouse for linear algebra algorithms.
Why do you think that's wrong? I simply claimed MATLAB is much more commonly used by scientists and non-SW engineers. Most SWE jobs are in B2B, web development, embedded, etc fields. Which one of these sees a lot of MATLAB usage?
Maybe don't resort to insults straight away too.
Citation?
> embedded, etc fields. Which one of these sees a lot of MATLAB usage?
Both embedded and etc. MATLAB is used widely in signal processing and controls software.
> Maybe don't resort to insults straight away too.
Sorry. I get annoyed when someone pretends that Software==Websites as if any of our trusted electronics are designed or programmed in javascript.
The main contrast Dijkstra is drawing is between closed and half open intervals. Typically, the choice "starting from one" is tied to "closed" intervals, but that's not strictly necessary. Nothing stops you from using half-open intervals on Lua lists any more than you can used closed intervals to index into C arrays.
The main advantage of the typical coupling is avoiding `+`s and `-`s in your "typical" loops over the first n indices:
for i = 1, n do -- closed
for (int i = 0; i < n; i++) { // half open
but what if you want to go in reverse? Getting the _final_ indexes requires arithmetic with half-open intervals, but it doesn't with closed (which also requires changing the operator from exclude-eq to include-eq): for i = n, 1, -1 do -- closed
for (int i = n - 1; i >= 0; i--) { // half open
Of course, probably the most important property of half-open intervals is the way that they can be broken down into two disjoint intervals: `[A, C) = [A, B) U [B, C)`. But if your indexes are integers this is straightforward with closed intervals too: `[A, C] = [A, B - 1] U [B, C]` or `[A, C] = [A, B] U [B + 1, C]`.But again, if you're using this property, you probably don't actually care about the "absolute" indexes at all; indexing with `[1, #list + 1)` isn't a problem _at all_ in Lua.
Plus, between the operations of "get final index" and "split interval into non-overlapping pieces", "get final index" is far more common (it's crucial to stacks, queues, etc, in addition to various uses like iterating backwards). Choosing "closed" intervals is then the cleaner option to avoid an ugly "+" or "-"!
Dijkstra mentions that requiring -1 to refer to an initial interval as "ugly" since -1 is not a valid index, but he has no qualms with using the length of the list, n, and allowing it to not be an index!
This is all said without even going into the awkwardness of having [k] being the {k+1}th element (e.g., [4] is how we write "fifth" in C -- I distinctly remember at least one conversation at work where the wrong thing was communicated because someone said "fourth" to mean [4])
Now, don't take this as a serious argument that 0-indexed-lists are inferior to 1-indexed lists. It's just a demonstration that it's easy to argue either side. That's because it's really just an unimportant convention. Pick one and stick with it and everything will work fine, modulo confusion about what "fourth" means.
That's not what string.sub accepts though.
> Plus, between the operations of "get final index" and "split interval into non-overlapping pieces", "get final index" is far more common (it's crucial to stacks, queues, etc, in addition to various uses like iterating backwards). Choosing "closed" intervals is then the cleaner option to avoid an ugly "+" or "-"!
It (rather unnecessarily) takes log time, so you probably shouldn't do it if you can avoid it.
This is one of the things that improved in Lua 5.4. Now it caches the table length so that the # operator is O(1) in many common situations.
That is, does #({[1]="a", [3]="b"}) return 1 or 2?
edit: I see that it still returns 1. That's a shame. This trips up so many newcomers who've been told that "#" means "length". IMO they need to either start calling it something other than the "length operator" or make it _actually_ return the length. Sadly at this point it would just be yet another breaking change.
edit: yes, in fantasy-lua there will be an array type that can represent [1,nil,3] (sequence of 3 values)
Some early alpha versions of Lua 5.4 played around with this idea but it turns out that it would break too many things. I'd expect that if Lua ever makes a major jump to Lua 6.0 then nils in tables will likely be one of the main changes.
On the other hand, my not-very-well-considered preferred approach creates some asymmetry between arrays and maps as to what values can be stored. Or it perpetuates the existing asymmetry between varargs and tables...
> t[i]=undef is a special syntax form that means "remove the key 'i' from table 't'". You can only use 'undef' in three specific forms:
t[i] = undef -- remove a key from a table
t[i] == undef -- test whether a table has a key
t[i] ~= undef -- test whether a table has a key > for (int i = n - 1; i >= 0; i--) { // half open
Well, there's your problem right there[0]; it should be: for(size_t i = n; i-- > 0 ;) {
> This is all said without even going into the awkwardness of having [k] being the {k+1}th elementYou mean the awkwardness of humans using "kth" to refer to element number k-1?
0: /sarcasm
To mark a node at the third position, containing one character, we say {3,3}.
But what if there's a node at the third position, that contains no character? Well, thats {3,2}.
It's like nails on a chalkboard every time. 1-based indexing is a mistake.
~=, by contrast, is just notation. You get used to it quickly.
Why not shift over indexing by one consistently? So you'd use {3, 4} and {3, 3} for each, so the length property is preserved.
("12345"):sub(3,3) returns "3", while ("12345"):sub(3,4) returns "34", and ("12345"):sub(3,2) returns "".
Indexing in the array portion of tables, and strings, is at least consistent; but Djikstra was right.
The Wiki[0] has a good discussion of Djikstra's iconic essay. The best of these arguments, imho: 0 based indexing unifies enumeration and measurement.
FWIW, I think that the real mistake was to make inclusivity/exclusivity implicit to begin with. If the syntax is explicit, and the choice is captured in the resulting value (i.e. it's more than just a pair of numbers), then you just use whatever is more convenient for the task at hand. Nim almost gets there: (x .. y) is end-inclusive, and (x ..< y) is end-exclusive - but the syntax still exhibits a preference.
Swift is similar, using `...` and `..<`. Rust uses an alternative convention, `..=` and `..`.
Would you prefer to see `..=` and `..<` so that inclusivity/exclusivity is always explicit?
x = calculate_n_dogs()
# later
ok = all(x.status_code == 200 for x in my_other_responses)
# later
let_the_dogs_out(x)
# oh no, why isn't x n_dogs anymorePython 2 did have the issue where list comprehensions (though still not generator expressions) didn't introduce a new scope, and so you could indeed get issues quite similar to your example above. This was corrected in Python 3.
All that said, I will agree that Python's scoping rules are pants-on-head crazy. In very nearly every other programming language under the sun, you declare variables at the point where they exist. In Python, you declare a variable when it exists somewhere else (and you then want to reassign it). This does mean that you don't need to put a "local" or "var" or what-have-you in front of each new variable you declare (which is basically the reason Python did it this way), but the semantics of the thing can take some getting used to.
I've run into similar issues where the variable was used in a for loop declaration rather than a list comprehension, but these are a bit easier to spot. Also, maybe people who don't use lua constantly won't be in the habit of thinking they can safely shadow locals with their looping variables.
It's a little unconventional, but then again ~ is a fairly standard (bitwise) negation operator.
> arrays start at 1
Lua doesn't do arrays, it does tables. Though they have similar syntax in certain cases, and are occasionally used for the same things, they are fundamentally different things.
[0]: where "array" means contiguous memory.
[1]: where "list" refers to a high-level collection of elements
I love the platform and ideals and goals - a simple, lightweight embedding language... but the language details rub me the wrong way.
- some syntax like js arrow function (this is tough because {} are not block delimiters, maybe single-expression lambdas like in python? I kind of hate those though.)
- `??` operator which does the same thing as `or` except when the left operand is exactly `false`
- `?:` operator. Should this evaluate to exactly one value or any number of values? e.g. x?2,3:4,5
I recently encountered a library that addresses some of the first need: https://github.com/starwing/luaiter#the-selector-interface
Theoretically I could get this stuff by using Moonscript, but then I'm using a somewhat esoteric language that brings in a lot of other very unusual ideas.
Although for ternary, I honestly think the "if/elseif/else can be used as an expression" has always been my favorite syntax. Or the CASE...WHEN syntax of SQL.
`or` does this unless there is a possibility your function returns boolean `false`
But honestly, I can't recommend Lua for any other use case. You'll be chugging along and then run into a brick wall when you need e.g. real regexes instead of Lua's comparatively crippled pattern matching, or a networking library that isn't hosted on one professor's personal web page that hasn't been updated in years. Take it from someone who had to throw out and port a bunch of Lua scripts when they just couldn't keep up with new business requirements: Python or Perl or Ruby is almost always a better choice.
It was. At peak perhaps 10 years ago. It is telling that the game engine citation in that article is from 2009.
Now the ecosystem of game engines is very different. Very few games are written in raw C/++ and have to have their own tiny scripting language packaged. They are mostly based on more established engines. 'Small studios' have gone, niched out of end-to-end dev, or re-invented as the 'indie game' world, and pretty much none have their own multi-game engine.
In fact, I've only come across Lua used in Pico-8 gamejam games in the last 2 yrs, personally. (Though I'm retired now.)
It was a phenomenal language for gamedev. I commissioned a Lua consultant to extend the language in a small way for a game, I ended up paying for about half the hours I expected.
Yeah, I wondered if that was still the case. Most of the examples I can find of “games that use Lua” are indeed around 10 years old.
Are more recent games using different scripting languages, or are they designed to not use a scripting language at all?
A lot of game devs use frameworks or engines that strongly encourage or require the user to use C#, javascript, gdscript, or gml. Solar2d and love2d use Lua but don't seem very popular. Fewer people are throwing together their own thing in C++ and putting Lua inside.
It's also used in Redis, and there's at least one major network router that uses it.
Lua is designed specifically to be embedded within a larger program to manipulate only the features of that specific program. It is for scripting your program. Not, for scripting-up a program in Lua.
Trying to do the opposite, embed Python as a scripting language inside of a larger program, is certainly possible. But, it is much, much larger endeavor that almost certainly drags in 100 pieces of functionality that you do not want inside your program for every 1 that you actually want to use.
Sounds like for your situation Python/Ruby/Perl are much better matches.
https://docs.racket-lang.org/inside/embedding.html
I've heard that some of the new WebAssembly implementations are actually nice to embed. But, that just moves the problem to "What language to compile to WebAssembly?" ;)
lpeg is leaps and bounds beyond regex. When I have to work with mere regex, I miss lpeg, the same way someone who is used to regex would feel when stuck with Lua patterns.
Networking isn't a problem for me personally, I just use luv and I'm very happy with it. But it's true that this is under-documented, and you do have to do more thinking and less Stack Overflow copypasting than you would with a language with a larger user base.
- https://docs.fluentbit.io/manual/pipeline/filters/lua
I am happy to see performance is being improved, as of now we stick to LuaJIT, since our use case is pretty simple we have not found a good reason to move away from it.
const x = 1
local const y = 2
over what they're adding.
local f <close>, err = io.open("foo.txt")
if not f then error(err) end
vs local <close> f, err = io.open("foo.txt")
if not f then error(err) end
In the first version it is clearer that only the "f" variable is marked as a to-be-closed variable. The second version looks nicer when there is a single variable but is more confusing when there are multiple variables. local f close = io.open("foo.txt")
seems fine. local f
close = io.open("foo.txt")
Despite the lack of semicolons, Lua is a free-form language and newlines are treated the same as any other whitespace character.A) these are properties of the variable, not the value. <close> triggers when the variable goes out of scope, and <const> supposedly was a feature that came out of implementing <close>. Both only work with local variables.
B) the angle brackets leave room for extension with more properties in the future without creating an additional new syntax.
C) you can multi-declare several variables with different properties i.e.: `local a <const>, b, c <const> = "apple", "banana", "carrot"` with only a and c being const.
D) no new keyword(s) added to the language
Compared to the simple syntax of the rest of the language it may be a little jarring, but it was well thought out and imo it's still better than weird stuff like "spaceship operators", "hashrockets" or other weird syntax oddities in more popular languages. Would you really prefer chains of keywords? `static const char *strprbrk` We have angle brackets now. If it turns out to be a mistake, I'm sure Roberto and his team will just take it back out of the language.
The new `const` and RAII/resource scope language features remind me of the golang discussion yesterday, someone linked Pike saying JS/TS, C++, Hack, etc keep borrowing features from each other. Of course we won’t fully converge on some perfect syntax any time soon, or ever, but we seem to be settling into a general consensus on syntax in general, and semantics of dynamic languages.