[1] Some did recognize this difficulty and introduced an additional pair of grouping characters (e.g. `%|asdf|` or `${asdf}`). These were however rare, and those that ultimately allowed an arbitrary expression were even rarer.
If I remember correctly these kinds of format strings have to be resolved by the compiler since they refer to variables in the source by name, which is information most languages do not keep around at runtime and a security nightmare if done dynamically. Also performance sensitive language like C aren't exactly known for a fast strcat, it is usually left as an exercise to the user to avoid catastrophic runtime behavior caused by the eternal search for nul.
(format "{a} {b} cde {f}")
Could be a macro that returns: (concat a " " b " cde " f)
The non-trivial part seems to be replacing the original format call site with the compiled version in the code. When I try to think it through, I conclude that everything should be a macro that outputs a compiled program. Just seems wrong in a dynamic language but it seems people do it all the time:https://news.ycombinator.com/item?id=39240528
Partially evaluated programs.
I'm not sure what you mean by "runtime introspection" in this context, because literals with embedded expressions don't require any runtime parsing or lookup - such a literal is parsed completely at compile-time, and in languages with lexical scoping, all variables etc are resolved then. It really is just syntactic sugar for concatenating; it's not at all like runtime templates.
Format strings live in userland as part of the standard library and have the many drawbacks you note.
Proper string interpolation is part of the language grammar itself, and the bits between the braces are usually full-fledged expressions that resolve variable names the same way that every other expression does. Since it's a language construct, it's trivial for it to just compile down to a series of concat(a, b) operations.
It sounds like Python had a half-complete string interpolation mechanism until this PEP, so I can understand why there'd be confusion on this point, but that is a problem with Python's implementation, not something that is an inherent risk with string interpolation. In a new language it is easier to implement string interpolation properly than to introduce a half measure like Python's.
xs = [2, 3, 4]
ys = [1, *xs, 5]
the old way feels backwards: xs = [2, 3, 4]
ys = [1] + xs + [5] ys = [1, ...xs, 5]
It is more visually indicative of a 'spread' of values.But overloading + for lists in general is a footgun anyway, especially in a dynamically typed language where the actual behavior is determined at runtime. And Python went one step further and also defined += for lists, and did it in a way that is even worse, because this:
xs += [1]
does not behave the same as this: xs = xs + [1]
(the latter creates a new list and assigns the reference to it to xs; the former mutates the list in-place without changing the value of the variable)Then on top of that you have scoping issues, too, because += is considered an assignment operator, and thus variable on the left side of it is considered a local unless you have a "global" declaration. So if you do this:
xs = []
def foo(): xs += [1]
it doesn't work as expected, because xs inside foo is a new local that shadows the global xs - and it fails at runtime because local xs is not initialized. You have to either do "global xs", or else write xs.append(1) in this context. At which point you might as well just use .append() everywhere, since at least it works consistently.All in all, it's a good illustration of what happens when people get too "clever" with syntactic sugar when designing a PL.
In the end interpolation is either dumb and insufficient for most advanced use cases or it is a complex mini language within the real language with tons of corner cases. Doing everything with language facilities looks much more elegant to me.
[1] The various forms of the concatenation operator (+, ., ||, &, :) are most common but not unique. Perl has a sort of scalar string multiplication where `"*" x 80` turns into 80 asterisks.
Python just lets you multiply strings, so you can do ("foo" * 80) etc. Which of course has all the same problems as overloading + for string ops. Perl had a better idea there, but that in turn is kinda forced by the design that implicitly converts strings to numbers and vice versa - you really, really don't want to be in a situation where 1+2, "1"+2, 1+"2", and "1"+"2" are all valid but inconsistent (looking at you, JS...).
Format strings are what you describe: either too weak or a completely insane sub-language.
The kind of string interpolation that OP is talking about is different in that what is in between the braces is just an expression and can usually be an arbitrary expression. It doesn't need to be a complex mini language because it just uses the same expression language that you're using everywhere else, with the only additional rule being that your expression must return a string (or in some languages something that can be coerced to a string).
From reading this PEP, it sounds like Python's original implementation was a bizarre half measure—it was almost like proper string interpolation but had a bunch of restrictions that showed that they stopped short of changing the parser to just go into expression mode in between the braces.