New string formatting in Python
zerokspot.com
zerokspot.com
What's more, it's being hailed at a plus for localization which it isn't. Localizers should never, ever deal with string interpolation - anything past what .format() does is essentially untranslatable.
That's because some languages have complicated changes in the text depending on eg the number of things: not just singular/plural, but more complicated. Russian is one example.
In practice, what you need is a notion of a "context" for your formatting (often something that is aware of both the current i18n and l10n factors), and that context needs access to the fundamental objects that would be rendered in to a string.
Effectively, the "format string" becomes an identifier for the particular message you want to render, though often the objects themselves are definitive enough.
In that paradigm, interpolation/formatting/whatever might be the default mechanism employed by the context, but you don't want to have that as your explicit mechanism. You want one more level of indirection before you get to it.
Some platforms make the mistake of having that context be tied to the thread, out worse still, a global, and that falls apart once you have any multiplexing logic (and the way python currently works with posix locales fully captures how terribly you can do this). Either way, I'd argue out really is a different problem from interpolation, and if you are trying to solve it using interpolation, you don't understand the problem.
I generally localize for western languages and write in many programming languages. The combination of gettext + python format strings has been working really great for me, and generally much better than other systems I've seen and put to use. In fact, the simplicity of gettext provides a very fast turn-around, and with translators experienced with the tool I never had problems. Python format strings also work great in this context, as I can supply an arbitrary dictionary of elements that the translator might need. The only real problem has been plural forms in complex text strings, where ngettext is not always sufficient.
What method (name a project I can inspect) do you recommend as a good localization architecture?
l8n_context.format('file_not_found', file)
'file_not_found' is just an identifier to lookup the actual format template (likely a format string) that will be combined with the file object to render the error message.(I think we are already agreeing, only that you express your point differently in Python specific terms. I don't care whether you stuff your code into a context object or something else. I was talking about having the full power of the programming language available, vs using a limited language like format strings.)
If translators are cooperating with you, it's very easy to provide the needed elements directly in the format's dictionary (that is: you extract the translatable pieces for them). It also means you don't have to worry that they're going to fiddle with mutable state.
I generally write everything in english, and do back-translation to my own locale (I also cooperate for translating external projects into my locale), so I eat my own dogfood here.
I know I do not want to deal with extra lower-level subtleties here. Translation is hard already by itself. It's impressive how a good translation of a simple UI can take so much time. If I had to inspect the object to know what I can get out of it I would get crazy.
I'd take a pre-baked dictionary any time.
I've also already used the string-catalog approach in the past (heh, XUL), and I'd personally take gettext any day.
At some point the translator will have to format some string himself.
I am not sure why i18n is a big deal, let another library deal with it.
Whatever, I have the choice not using this feature after it is accepted and implemented. Their PEP discussion on email always ended up in tangential. Problem I always have with string manipulation is dealing with long string, which for coding style I'd split into multiple +, and thus using format is pretty ugly.
>>> (" foo "
... "bar.")
' foo bar.'That was annoying.
It also has little in common with i18n, the use cases differ too much. Perhaps in the future someone can figure out how to bring them together, but not today.
Go and Rust are two notable new languages which require explicit formatting functions as opposed to relying on special string interpolation syntax.
Is it really worth updating all the Python syntax formatting and analysis code out there just to save one character on an operator? I don't think it's a good tradeoff.
You got:
price = 18.8
date = datetime.datetime.now()
Now you want to save in a file: "It's 18.80 and we are the 01/04/2016"
With +: num, dec = str(price).split('.')
msg = ("It's "+ num + "." + dec.ljust(2, '0') +
" and we are the " + date.strftime("%m/%d/%Y") + "\n")
f.write(msg)
It is so long and ugly we have to break it on several lines to respect PEP8.And it's a pain to write, or even worst, to read.
With print():
num, dec = str(price).split('.')
print("It's", num, ".", dec.ljust(2, '0'),
" and we are the ", date.strftime("%m/%d/%Y")), file=f)
A tiny bit better since we don't have to use that many concatenation tricks.If we didn't have to format the float, print() would have converted it to string for us which is nice (with + you have to call str() all the time).
With "%":
f.write("It's %.2f and we are the %s\n" % (price, date.strftime("%m/%d/%Y")))
Way better, but still hard to read.With format():
f.write("It's {:.2f} and we are the {:%m/%d/%Y}\n".format(price, date))
Now we are getting sometwhere...With fstrings:
f.write("It's {price:.2f} and we are the {date:%m/%d/%Y}\n")
fstrings have everything: - easier / faster to type than any other versions;
- easier to read than any other versions;
- as expressive as any other versions.
It's reduce work by making you think less, type faster, read faster (for you and your colleague) and it's easier to spot bugs.The only place where you don't want fstring are for l10n.
For example, your i18n tool can extract:
"It's {price:.2f} and we are the {date:%m/%d/%Y}\n"
as a key, and then l10n can map this to: fr_FR: "C'est {price:.2f} et nous sommes le {date:%d/%m/%Y}\n"
(Here we can fix up the silly US date formatting, too, although really you shouldn't be localising dates in your format strings).All that remains is for them to work out some way of mapping in that localised string, as you can do this:
_("It's {price:.2f} and we are the {date:%m/%d/%Y}\n").format(price=price, date=date)
any more.ISO 8601 (similar to Japanese format) is the most logical way to format dates: YYYY-MM-DD, easily sortable, not the silly way US and Europeans format their dates. ;)
If you are a fan of ISO8601 (which I am), you can then set this globally on your desktop, rather than expect the app developers / translators to choose for you.
>>> '%(one)s %(two)s'%{"one":"hello", "two":"world"}
'hello world'
>>> '{one} {two}'.format(one="hello", two="world")
'hello world'
The PEP that is looking at this problem has been deferred: at the end of the document they kind of indicate that they didn't really think about the i18n problem correctly and so don't actually have a good solution to present, and are going back to think about it more.Yes, having so many solutions is not ideal, espacially when it comes to teach the language. However, it would be foolish to avoid improving the usability of Python just to avoir having "one more way to do it".
I do wish they'd deprecate Template though. It's more than useless.
"{a} {b} {a}".format(a=a, b=b)
"{a} {b} {a}".format(**locals())
Compared to this: f"{a} {b} {a}"
Sorry that's about 1000% better. This should have been the one way to do it, originally. It isn't magic either, rather a simple compile-time transformation to existing format syntax. There's nothing new to remember besides a large reduction in noise. with open('template.txt') as f:
template = f.read()
formatted = template.format(**values)"PEP 498 proposes new syntactic support for string interpolation that is transparent to the compiler, allow name references from the interpolation operation full access to containing namespaces (as with any other expression), rather than being limited to explicit name references. These are referred to in the PEP as "f-strings" (a mnemonic for "formatted strings")."
"Full access to containing namespaces?" From strings? Bad, bad idea. This is currently marked as "deferred", but should be marked "rejected with extreme prejudice".
myquery = sql(i"SELECT {column} FROM {table};")
What could possibly go wrong? myquery = sql("SELECT %s FROM %s" % (column, table)) i"SELECT {settings.SECRET_KEY};"
PS: (I don't know why but HN is not updating the page with the reply links needed, so I'll just edit this)Yes, I agree. I think that the misunderstanding happened when mixmastamyk wrote
> That's a good thing due to security reasons as arbitrary expressions are allowed. There are plenty of templating solutions available
The point is that even if you don't allow arbitrary expressions (which imho are a mistake, and of which I haven't seen a single use case yet), having this kind of interpolation from strings that are not literals (i.e. are not in the trusted source code) would still be a security issue
Since the PEPs apparently don't propose to extend this to non-literals, we're safe. But it's better to be wary and attentively review such proposals...
In fact, I just realized right now that Animats might have misunderstood PEP 501, since
sql(i"SELECT {column} FROM {table};")
should be perfectly safe from sqli vulnsPPS: Unless Animats is pointing out how switching i'' for f'' is a terribly simple mistake to do and hard to spot during a code review... I agree with that
"SELECT {};".format(settings.SECRET_KEY)
Remember that the interpolation only works for string literals, you can't inject that from external input. fmt("{a} {b} {c}", a=a, b=b, c=c)
If you want to save typing, maybe use :a instead of {a}. Or ?a would have made plain old ? a nice positional variant: fmt("?a ? ?", a=a, b, c)
The main benefit of the fmt function is that it requires no syntax changes to the language and is trivially provided by a third party library for all past versions of Python.That being said this ship has sailed. I guess I just take a more conservative approach to syntax changes than most.
Update: a bit sad to see my votes fluctuating wildly on this post. Please don't use votes to support or disagree with me: that's not what they're for. Please vote only based on whether you find this relevant.
I consider specifying the values or variables alongside the formatting string a requirement to be considered explicit, but I can see how it's a matter of opinion.
@decorator
def decorated
Is doing: decorated = decorator(decorated)
The switch is implicit too. It's syntaxic sugar to gain a pratical and elegant syntax for a common use case.It's not going to introduce vulnerability. It's going to make your code easier to write and read. It's going to make bug easier to spot in formating. It's going to make shell sessions easier.
I'm not trying to be pedantic. It really does affect readability when syntactic sugar's affect spans an entire scope. Whether or not that effect on readability is greater or less than the gain by the syntactic sugar is always a matter of opinion. Obviously my opinion is out of line with Python's core devs.
>>> f"foo"
File "<stdin>", line 1
f"foo"
^
SyntaxError: invalid syntaxThis is not true and exactly the distinction I'm trying to make:
Adding new packages, functions, objects, etc. can all be backported to older versions and alternative implementations. They also require no updates to ASTs, linters, syntax highlighters, static analysis tools etc.
Adding new syntax is backward incompatible (unless it's added as a from __future__ import to new old releases) and requires changes to all tools that parse Python syntax (the interpreter, ASTs, linters, transpilers, etc).
It is a shame that linters will have to add a letter to their grammar also, but I argue that the everyday usability and readability for millions will outweigh this drawback.
It isn't a large syntax change, string prefixes have existed since the beginning.
People forget that Python is also about readability, and ease of code exploration is the shell. Fstring is a feature to make that even better.
_(f"English {a} words {b} here {a}")
The interpolation is done before the string is passed to gettext which can't retrieve the translated string any more. "{} {}".format(a, b)
"%s %s" % (a, b,)
That's not very convincing. I was hoping this article would make a good case for interpolated strings, since it's starting to feel like Python is having an identity crisis. Type annotations especially took me by surprise, but string interpolation is another good example of an addition that doesn't feel like Python (imho, anyway). >>> very_long_var_name_1 = 'spam'
>>> very_long_var_name_2 = 'ham'
compare >>> # Explicit but tedious and doesn't help readability:
>>> print('{very_long_var_name_1}: {very_long_var_name_2}'.format(
... very_long_var_name_1=very_long_var_name_1,
... very_long_var_name_2=very_long_var_name_2))
spam: ham
with >>> # Explicit but somehow feels dirty:
>>> print('{very_long_var_name_1}: {very_long_var_name_2}'.format(
... **locals()))
spam: ham
and >>> # Still fits on one line. I think f prefix makes intent clear.
>>> print(f'{very_long_var_name_1}: {very_long_var_name_2}')
spam: hamThe best way imho would be:
vars = {'short1': very_long_var_name1,
'short2': very_long_var_name2}
print('{short1} {short2}'.format(**vars))
Easy and extremely unlikely to ever include the wrong variable.That's false, it does not do that. It converts the string into the equivalent format call at compile-time. RTFP ;)
e.g the following is an error.
a = 4
"a: " + a
You have to do: "a: " + str(a)A minor correction to your comment: "+" calls the __add__ method (big surprise) - just fyi
>>> from string import Template
>>> s = Template('$who likes $what')
>>> s.substitute(who='tim', what='kung pao')
This will probably be TIL for many people. It is a surprisingly hidden feature.I for one, still like the "%s" % x instead of "{0}".format(x). It is simply shorter and I already know or use printf for C in other parts of the project. But with the new f"..." interpolation, I can see liking that more. I am all for being as concise as possible while still being explicit (if that makes any sense at all ;-) ).
In the same way, .format() solved the problem of inconsistency and bugs happening all the time with %. It was a restriction and formalization effort, trying to root out bad practices and hence slightly more explicit, but it made sense.
The new f'' IMHO does not solve anything beyond pandering to developer laziness. It will likely introduce bugs in places where developers are not clear about the context ("oh, I thought we didn't have a 'x' var at this point, turns out we do!"), because (from what I understand) it takes away the ability to define which variables should be considered.
Luckily it will take a while before it percolates in any significant library, but I sincerely hope it just doesn't gain much traction.
"There should be one-- and preferably only one --obvious way to do it."
There are a lot of warts on that snake, but I still really like using Python.
#python
import re
m = re.search('(a.+)(d.+)', 'abcdef')
if m:
print(f"{m.group(1)},{m.group(2)}")
#perl
if('abcdef' =~ /(a.+)(d.+)/){
print "$1,$2";
}That aside, though, I don't consider brevity alone the most critical criteria for a programming language; that way lies APL. Expressiveness, yes, but not at the expense of clarity.
Also "Simple is better than complex", "practicality beats purity" and "Readability counts".
The zen is not the bible, you don't get to cherrypick the stuff you want to make your case.
Plus, they are just guide lines, in the end, you have a debate in the python dev mailing list with reasonable people making their case.
"%s %s %s" % (a, b, a)
"{} {} {}".format(a, b, a)
If you want placeholders to match variable names, you can do: "{a} {b} {a}".format(a=a, b=b)
"{a} {b} {a}".format(**locals())
So this is just unnecessary (especially since the "f" prefix is easier to miss than a "format" method): f"{a} {b} {a}"
And this is downright obfuscated—putting operators inside of string literals: f"{a + ' ' + b + ' ' + a}"https://www.python.org/dev/peps/pep-0498/#no-use-of-globals-...
You could argue that's not a good enough reason, but it's there.
You've all been programming in C-like languages for far too long to realize what a horrible design string formatting is. You can argue over "explicitness" all you want, the new way is easier to learn, easier to read, makes more intuitive sense, requires learning fewer rules, ad is close enough to the format string method that they work well together.
The one counter argument that makes sense to me is that in general we shouldn't be doing easy string interpolation, since that way lies SQL injection, XSS, etc, and should instead rely on a stronger type system with binary text blobs, HtmlStrings, SqlStrings, etc, with automatic escaping into and out of the data type.
But then that's not the case with Python now. If you're only trying to stick this string inside that string in a quick and dirty manner, I totally don't understand the reticence folks have to something the way ruby does it: "Name: #{first_name}".
Don't get me wrong, I'd like Python to have better, more obvious, more concise string formatting. However the last time we had this discussion, it was about str.format() and how it was going to be awesome and don't worry modulo-formatting will go away.
Turns out it did not; modulo formatting is still there because why would it be removed. This is history repeating itself - are you actually baffled that some people learn from past mistakes?
None of the other techniques are going away. This time yes, no one is naive enough to think so.
As a ruby dev posted here, it is obviously better in most respects in most common cases.
It may be true that
"{} {}".format(a, b)
is a bit verbose, but it is crystal clear and clean. Just remember the Python Zen: "There should be one-- and preferably only one --obvious way to do it." _and_ "Explicit is better than implicit."Approximately 2.5 times since then has whitespace been a problem, and which I fixed in under 10 seconds each time. Yet, the readability gains from removal of block delimiters in that time frame is uncountable.
> Saving a few keystrokes writing ".format" doesn't even register on the same scale of annoyance
It isn't just .format, it is:
.format(long_variable_name1=long_variable_name1,
long_variable_name2=long_variable_name2)
This is a huge win that should have happened long ago.from __future__ import string_interpolation
:)
...Or just do it in 28 lines of Lua:
http://hisham.hm/2016/01/04/string-interpolation-in-lua/
This is a nice showcase of how Lua's metamechanisms can be applied to do things that often require new features in other languages.
Now if python would begin to support immutable values by default then I'd be most content, and Python complete enough.
Hoping for a decision from the core team – having both f"" and .format is a pretty clear deviation from this principle.
If Python 3.6 is going to introduce multiple ways to do the same thing, there is no good reason to not merge Python 2 and 3 together and have both set of behavior co-exist with each other (__future__ or __past__).
In one shot, you break the Berlin wall of Python.
- class stuff(object) vs class stuff;
- range vs xrange;
- itertools.izip vs zip;
- itertools.imap vs map;
- itertools.ifilter vs filter;
- dict.items vs dict.iteritems vs dict.viewitems;
- dict.items vs dict.itervalues vs dict.viewvalues;
- dict.items vs dict.iterkeys vs dict.viewkeys;
- __cmp__ vs __eq__ + __gt__;
- sorted(cmp) vs sorted(key);
They didn't fallout from the philosophy. They have been pragmatic and tried to balance the language design : gaining modern features vs making a robust base vs pleasing the legacy crowd. Is. Is. Very. Hard.And yes, we would all prefer to have less way to format. Would it mean I would prefer NOT to have fstring ? Certainly not, it's a great feature. We can't live in the past because it will make us not stand perfectly to the ideal we have.
Real life is not ideal.
But merging Python 2 and 3 ?
With the string model completly reworked, that would be apocaliptic. Most people don't realize how deep the unicode change has been.
I have been dev and teaching Python 2 and 3 for years. The amount of problems linked to UnicodeDecodeError dropped by 90% after the switch.
Not because Python 2 model didn't work.
Because nobody understand text.
Most dev don't understand what text is. They just want to format string. That's what Python 3 helps to do, and it does it well.
Mixing both would be like mixing olive oil and vanilla ice. Great on their own way, but use them together all you'll get a terrible meal.
Do you seriously think anybody is thinking of dropping python 2 support by 2020 ? that will only create a fork. None of the core frameworks have upgraded in a decade. Look at Flask for example. So yes, I havent been using Python 3 much - there has been no reason to.
the point I'm trying to make is how to get everyone on the same page. The reason Python 3 was api incompatible was because of the core tenet of one-true-way. All the functions you mentioned may be superior, but unless you give people a way to mix and match both in the same source , you will not have adoption.
Or do you think Python 3 adoption has been successful ?
P3 '.format 'is fine, the only problem I have is forgetting the last ')' and vim picks this up. Is interpolation that good to introduce another way of doing things?
So funny ... the most prominent language known for this stuff is PHP ... and missing :D
That statement is against my religious beliefs. Apparently we can stick an f in front of the string, but we can't have default values that work with locals()?
Adding a new syntax that doesn't exist in any language.
vs.
Using sensible default parameters.
-----------------
An argument for why my suggestion would be bad: https://www.python.org/dev/peps/pep-0498/#no-use-of-globals-...
Not true, look into the new interpolation features in C#, Scala, JS, Swift, nim, etc.
Please, even JavaScript (ES6) got itself something similar; implicit string interpolation is not that rare.
Also, this alternative solution would not allow arbitrary expressions, which was a design goal.
f(**{x:x for x in d})Well, star-star locals. I don't know how to write it according to https://news.ycombinator.com/formatdoc