My experience binding a couple of scripting engines with C++
germandiagogomez.medium.com
germandiagogomez.medium.com
I can fully understand why lua is not a good fit for this case, however, I would like to add some color to the picture.
The most powerful way to for C(++) - lua interop, is not the official CAPI but luajit/FFI: https://luajit.org/ext_ffi.html
This allows for allocation of C objects on the heap and on the stack and FAST function calls. Doing the same for C++ is possible but requires some work e.g. http://lua-users.org/lists/lua-l/2011-07/msg00492.html
Furthermore:
- unfamiliar syntax -- The syntax is tiny -- and I found nothing unexpected about it.
- no classes -- There are many class libraries available for lua. Just pick one. Used penlight classes quite a bit, without running into major issues.
- use tables for both hash tables and arrays. -- Yes. This is on the API side, under the hood hashes and arrays are used where appropriate.
Alternatively, embrace the prototype-based programming and just... don't add classes.
> unfamiliar semantics: no classes, use tables, start indexing at 1 and other oddities, just as being able to call functions with the wrong number of arguments and returning nil on the way
Other than 1-based indexing, those semantics should be very familiar from JavaScript.
But it does not fullfill well the zero-friction I was looking for when integrating. You have to change a bit your mindset when integrating and using this stuff.
In Wren you have lists, which are basically arrays and dictionaries, ranges and all things I already know how to use from Python/C++.
That said, inheritance is not working smooth and ChaiScript does not even have inheritance itself. But for my purposes a class, concurrency and familiar data structures and patters was enough.
I've spent the last twenty-ish years in C and string manipulation just sucks. Think about it, you declare a buffer of length 20, the indexes are from 0->19, and the 19th byte needs to be a null if you are using it as a string and are using the entire buffer. Also, the standard library is not guaranteed to null terminate in all situations.
Lua's string indexing feels far more natural to me.
The mismatch with the English language and how people naturally count is definitely there and annoying. And yes, string manipulation in C is especially broken although I think the indexing is the smallest problem there.
However, it’s extremely natural when you think about it in terms of memory access. For example, in a 1-based indexing system, ptr[0] would point 1 character behind your pointer (weird) and ptr[-1] would point 2 back (wtf). Having the index map neatly to the offset makes a lot of sense to me. In fact, when I first started programming in VB6 20 years ago and only had a math background, the 1-based indexing was natural but I could never figure out why I had so many bugs related to array and string offsets.
I’ll also note that most programming languages are 0-based and interop with C is not really the goal (Java, JavaScript, Ruby, Python, etc). In fact, Python and Perl’s string manipulation is some of the best out there and they are 0 indexed.
a b c d
0 1 2 3 4
From this 4 character string you can slice s[1:3] and get 'b', 'c'. s[i:j] will always have j - i characters (ignoring negative indices). s[i] is as if it were short for s[i:i+1] e.g. s[3] is s[3:4] and that gives you 'd'.Admittedly C++ doesn't have slicing so this is less relevant, but I think it's still an interesting aspect to the discussion (without getting bogged down in ideology).
I completely agree regarding memory access, but would argue that strings and memory should not be treated in the same way. Having the index represent the length - 1 has caused countless off-by-one bugs [1] that would not have been there in the first place if string indexes started with 1.
Java, JavaScript, Python, (maybe Ruby, I'm not fluent there) also bite the user if you attempt index data outside of the string's range. C will happily index whatever you want, and these bugs can often remain hidden for decades.
Sure, 17-7 gives you 10, but that's not the final answer, you have to add 1 to get the right answer, 11. Sorry, no, you actually subtract 1 and the right answer is 9. Wait, no: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, okay, 11 in total, so you have to add 1, got it right the first time.
Lua is a great language for sure though.
Tuples are 1 indexed, lists are 1 indexed (in the standard library, anyway), but binaries are zero indexed.
Dijkstra: Why numbering should start at zero https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
The first element is at index 0 because its address is base + 0 * sizeof(element)
The second element is at index 1 because its address is base + 1 * sizeof(element)
1. Concurrency support
2. Easy to bind into C++
3. Familiarity, etc.
Namely, if only Lua had existed and provided that Sol2 exists and makes it bind it to C++ easy, I would have chosen that. But since Wren + WrenBind17 existed, Wren had more familiar syntax and it was a viable choice, I went for that. I was trying to find the past of least resistance (lower learning curve, easier to bind, concurrency making my code easier, since I am familiar already with most patterns)
As for ChaiScript, it was the first thing I took since it was so easy to embed. But it had its own problems: lack of concurrency and it does not point the file and line of errors, which is very painful because it drops your productivity.
And scripting... scripting is about productivity, at least that is what I was using it for.
Depending on your needs (small vs. fast, low memory usage vs. full JIT, etc.) you can pick anything from JerryScript / QuickJS all the way to V8.
Perhaps the only thing missing is a universal C/C++ API for embedding JavaScript engines that lets you swap out easily to test different trade-offs.
The ones you mentioned, as far as I investigated, were not dead-easy to integrate into C++, one of the top requirements.
Take into account that I have to expose my own types, not just ints and basic types.
The original API was coded in a natural C++ way. APIs that wrap well in that sense could be Sol2 for Lua, Chaiscript, Wrenbind... which integrate with custom types and smart pointers.
With other scripting languages and their libraries you need additional work
QuickJS’s API isn’t particularly well-documented, but it’s not hard to find your way around it either if you dig into the source (the engine is very hackable, too; you might even fix some of the language’s design flaws – obligatory wat talk reference – if you’re so inclined). The host API follows the CPython model, with objects represented by pointers and explicit reference counting on the C side. There are some predefined macros to ease defining built-in classes. Some type-level hackery in C++ might ease things even further. I don’t know how much deader-easier you want it.
These bindings do from decent to great.
> Lua ... It has unfamiliar syntax ... start indexing at 1
Syntax is not unfamiliar, just more Pascal like; if you use LuaJIT you can use zero based indices and a powerful FFI for direct C code integration.
Actually Lua is something to consider from the point of view of usage: it is an industry standard actually. However, all those small quirks in semantics... and classes can be done in many ways (that is what I understand, via metatables)...
In ChaiSCript or Wren there is one true way and you are done. You might like it or not, but it leads to less confusion, especially if you use most of the time what is in the mainstream.
This is by no means a bad thing in itself, it is just about how ergonomic or time-consuming it could be for myself: I just feel more comfortable with ChaiScript, Wren or Squirrel than with Lua. Even AngelsCript is also more similar to what you already have. So when exposing APIs there is much less friction.
Truth to be told, there is also https://github.com/ThePhD/sol2 which looks great and something to consider. It makes binding things quite easier and gives you object-oriented Lua. You could rely on that.
It was just my subjective choice. There is no 100% right choice. Probably, if I found people that are comfortable with Lua I would use that. But the case is that this is a project of mine as it stands now.
Not entirely, Lua standard libraries still expect everything to be one-indexed, while FFI structures are zero indexed. So with LuaJIT you often end up with a mix of 0 and 1 indexed code, which in my experience was workable but definitely a pain point.
> stuff = {"one", "two", "three"}
> stuff[1]
one
These native Lua structures are baked in, and they're a lot more flexible than FFI structures - presumably if you're using Lua, it's because you want to take advantage of that flexibility and those affordances. FWIW I've written a lot of LuaJIT, and usually I kept the lower-level FFI stuff separate from the higher-level code using Lua data structures, so I rarely encountered that discrepancy between them, but still something to keep in mind. > stuff = { [0]="one", [1]="two" }
> print(stuff[0])
one
Works well; I wrote e.g. https://github.com/rochus-keller/Smalltalk#a-smalltalk-80-in... that way.EDIT: even this works
> stuff = { [0]="one", "two", "three" }
> print(stuff[0]) -> one
> print(stuff[1]) -> twoIn this case, is it possible to make iteration start with the element at index 0? Maybe by implementing a custom version of `ipairs`?
> stuff = { [0]="one", "two", "three" }
> for k,v in pairs(stuff) do print(v) end
it prints all three elements in the correct order. I rarely use iterators for performance reasons anyway. Instead of ipairs one can use > for i=0,#stuff do print(stuff[i]) endhow is that not unfamiliar
No matter how you slice it or dice it, Pascal and Pascal-like syntax are not some obscure niche languages...
no, it is confined to people who learned Pascal ? I could start learning chinese with duolingo just as easily as I could start learning pascal with some web tutorial, it does not mean that either are familiar to me.
"Modula-3 is the origin of the syntax and semantics used for exceptions, and some other Python features." (from https://docs.python.org/3/faq/general.html#why-was-python-cr...). Also the predecessor languages ABC and SETL were in the Algol tradition.
> Python is 100000x more popular than Pascal
It's about factor 8 on https://www.tiobe.com/tiobe-index/ or factor 3 (in score) on https://spectrum.ieee.org/static/interactive-the-top-program..., whatever you prefer as a reference. Delphi (which is Object Pascal) is still a widely used language.
Pascal was discarded a few years back in my university. And yes, by familiar I mean exactly what you mean: you see nowadays Java, Python, C++, C, C#, but Pascal is disappeared.
Disappeared since long ago since I do not know even the syntax myself by casual reading around.
I'm wondering what you mean exactly by this. Do you not think about types when programming with a language with dynamic typing?
What I mean is that if you have to annotate all your code with types (like in AngelScript), this will slow you down for two reasons. First, you need to think about types, and second, refactoring is more rigid.
If it is optional, it is ok, you can take advantage of it at will (ChaiScript supports types in parameters, but optionally).
https://www.angelcode.com/angelscript/sdk/docs/manual/doc_da...
The Pyodide project has already compiled CPython to WebAssembly - why is that a worse solution than compiling one of these other scripting language interpreters to WASM?
On my i5 laptop, this demo downloads about 8Mb and takes a couple of seconds to load up: http://karay.me/truepyxel/demo.html
Lua, by comparison, is very small and has a fast startup under WASM.
Perhaps something like MicroPython could solve that, though.
With https://github.com/pybind/pybind11 there is really great integration with C++ and Python is my second home after C++ actually.
Anyway, I am quite happy with Wren and it seems to be fast (not a requirement for my project, though)
Only thing I can imagine is that it's not widely known and even more less documented.