PEP 701 – Syntactic formalization of f-strings
peps.python.org
peps.python.org
I've got some great news: https://cdn.zappy.app/339d2e16228a195e5149cdcf05edd572.png
Also one better: how to actually use it. There can be differences between the PEP and its exact implementation.
[1] Some did recognize this difficulty and introduced an additional pair of grouping characters (e.g. `%|asdf|` or `${asdf}`). These were however rare, and those that ultimately allowed an arbitrary expression were even rarer.
If I remember correctly these kinds of format strings have to be resolved by the compiler since they refer to variables in the source by name, which is information most languages do not keep around at runtime and a security nightmare if done dynamically. Also performance sensitive language like C aren't exactly known for a fast strcat, it is usually left as an exercise to the user to avoid catastrophic runtime behavior caused by the eternal search for nul.
(format "{a} {b} cde {f}")
Could be a macro that returns: (concat a " " b " cde " f)
The non-trivial part seems to be replacing the original format call site with the compiled version in the code. When I try to think it through, I conclude that everything should be a macro that outputs a compiled program. Just seems wrong in a dynamic language but it seems people do it all the time:https://news.ycombinator.com/item?id=39240528
Partially evaluated programs.
I'm not sure what you mean by "runtime introspection" in this context, because literals with embedded expressions don't require any runtime parsing or lookup - such a literal is parsed completely at compile-time, and in languages with lexical scoping, all variables etc are resolved then. It really is just syntactic sugar for concatenating; it's not at all like runtime templates.
Format strings live in userland as part of the standard library and have the many drawbacks you note.
Proper string interpolation is part of the language grammar itself, and the bits between the braces are usually full-fledged expressions that resolve variable names the same way that every other expression does. Since it's a language construct, it's trivial for it to just compile down to a series of concat(a, b) operations.
It sounds like Python had a half-complete string interpolation mechanism until this PEP, so I can understand why there'd be confusion on this point, but that is a problem with Python's implementation, not something that is an inherent risk with string interpolation. In a new language it is easier to implement string interpolation properly than to introduce a half measure like Python's.
xs = [2, 3, 4]
ys = [1, *xs, 5]
the old way feels backwards: xs = [2, 3, 4]
ys = [1] + xs + [5]But overloading + for lists in general is a footgun anyway, especially in a dynamically typed language where the actual behavior is determined at runtime. And Python went one step further and also defined += for lists, and did it in a way that is even worse, because this:
xs += [1]
does not behave the same as this: xs = xs + [1]
(the latter creates a new list and assigns the reference to it to xs; the former mutates the list in-place without changing the value of the variable)Then on top of that you have scoping issues, too, because += is considered an assignment operator, and thus variable on the left side of it is considered a local unless you have a "global" declaration. So if you do this:
xs = []
def foo(): xs += [1]
it doesn't work as expected, because xs inside foo is a new local that shadows the global xs - and it fails at runtime because local xs is not initialized. You have to either do "global xs", or else write xs.append(1) in this context. At which point you might as well just use .append() everywhere, since at least it works consistently.All in all, it's a good illustration of what happens when people get too "clever" with syntactic sugar when designing a PL.
ys = [1, ...xs, 5]
It is more visually indicative of a 'spread' of values.In the end interpolation is either dumb and insufficient for most advanced use cases or it is a complex mini language within the real language with tons of corner cases. Doing everything with language facilities looks much more elegant to me.
[1] The various forms of the concatenation operator (+, ., ||, &, :) are most common but not unique. Perl has a sort of scalar string multiplication where `"*" x 80` turns into 80 asterisks.
Format strings are what you describe: either too weak or a completely insane sub-language.
The kind of string interpolation that OP is talking about is different in that what is in between the braces is just an expression and can usually be an arbitrary expression. It doesn't need to be a complex mini language because it just uses the same expression language that you're using everywhere else, with the only additional rule being that your expression must return a string (or in some languages something that can be coerced to a string).
From reading this PEP, it sounds like Python's original implementation was a bizarre half measure—it was almost like proper string interpolation but had a bunch of restrictions that showed that they stopped short of changing the parser to just go into expression mode in between the braces.
Python just lets you multiply strings, so you can do ("foo" * 80) etc. Which of course has all the same problems as overloading + for string ops. Perl had a better idea there, but that in turn is kinda forced by the design that implicitly converts strings to numbers and vice versa - you really, really don't want to be in a situation where 1+2, "1"+2, 1+"2", and "1"+"2" are all valid but inconsistent (looking at you, JS...).
Especially allowing backslash and arbitrary quotes inside the expressions is a welcome and useful change, as this makes f-strings work more as expected without having to remember a lot of special edge cases and work arounds.
$ python3.12 -m timeit 'f"Welcome to {2 ** 3}."'
5000000 loops, best of 5: 69.8 nsec per loop
$ python3.12 -m timeit '"Welcome to {}.".format(2 ** 3)'
2000000 loops, best of 5: 112 nsec per loop
$ python3.12 -m timeit '"Welcome to %d." % 2 ** 3'
5000000 loops, best of 5: 82.2 nsec per loop
$ python3.12 -m timeit '"Welcome to " + str(2 ** 3) + "."'
5000000 loops, best of 5: 77.5 nsec per loopThe point he made was to embrace the new features of Python, and put old Python out of its misery, instead of wasting your time with backwards compatibility.
Just call whatever you're doing in the latest version of Python a "prototype", then ship it.
At 10:20 he talks about things you can do to put old Python out of its misery (dead parrot), including f-strings and ordered dictionaries.
The Fun of Reinvention (Screencast):
https://www.youtube.com/watch?v=js_0wjzuMfc
Invited Keynote Talk from PyCon Israel, June 12, 2017. I cause trouble and build a framework using all sorts of new Python 3.6+ features.
$ python3.12 -m timeit -s 'import string' -s 'x = string.printable * 50' 'f"Welcome to {x}."'
2000000 loops, best of 5: 106 nsec per loop
$ python3.12 -m timeit -s 'import string' -s 'x = string.printable * 50' '"Welcome to {}.".format(x)'
2000000 loops, best of 5: 200 nsec per loop
$ python3.12 -m timeit -s 'import string' -s 'x = string.printable * 50' '"Welcome to %s." % x'
2000000 loops, best of 5: 169 nsec per loop
$ python3.12 -m timeit -s 'import string' -s 'x = string.printable * 50' '"Welcome to " + str(x) + "."'
1000000 loops, best of 5: 214 nsec per loop
$ python3.12 -m timeit -s 'import string' -s 'x = string.printable * 50' '"Welcome to " + x + "."'
1000000 loops, best of 5: 198 nsec per loop
len(string.printable * 50) = 5000Never really understood why not all strings were f-strings in python.
For conditional formatting (where you don't know if the formatting is needed or the contents of what you're formatting into the string but you know what you want to format), f-strings also lose their use pretty much immediately.
They're great for a quick inline format but less useful for things like string templates you want to use over and over.
I like having different types of strings with different sets of magic characters. Even having identical `'` and `"` strings just to avoid backslashes in a few cases is great.
Now? Maybe because another breaking change in strings is the worst nightmare of the current generation of Python developers.
Sometimes you don't want interpolation, and it seems more reasonable to default to a version without such a feature than to default to a version with it and have an alternate syntax for non-interpolated strings. Or, at least, that makes sense to me, but I have no knowledge of their motivations.
If Python were designed from scratch today, I suspect it'd do something similar to JS.
# Ruby
"#{ "#{1+2}" }"
# JavaScript
`${`${1+2}`}`
# Swift
"\("\(1+2)")"
# C#
$"{$"{1+2}"}"
To which I might add # Perfect
"\{"\{1+2}"}"
basically Swift but { instead of ( as { is 'more decorative' so a better indicator that something special is going on. \ is already the escape character for such thing as \n and friends, so no extra escaping needed for e.g. ${I would continue to use .format, but all my teammates are converting everything to fstrings, and I think my cognitive load is much higher because of it.
Or were you thinking of anything else? I am not saying that a mutable string type should not be provided, just that it should not be the default one.