Ruby-like string interpolation in Python
github.com
github.com
>>> package = "foo"
>>> whatever = "bar"
>>> "Enjoy {package}".format(**locals())
'Enjoy foo'
(disclaimer: I wouldn't actually use locals() like this in real code)Also, if this thing is truly mangling bytecode, it's not portable between different python versions
But i'm not quite sure about that: I only skimmed the codebase, and that works seems to be done by interpy_untokenize, which boils down to some string mangling
Also, having expressions (or worse, statements... like it would be in ruby since there's no difference there) evaluated when evaluating a string is quite bad (this is not Haskell, and thus we cannot have guarantees that side effects won't happen)
Nice hack, btw
Consider the following Python2/3 code:
def outer():
x = 1
def inner():
print(x)
print("x = {x}".format(**locals()))
inner()
outer()
This actually prints the right thing when run: 1
x = 1
However, if you remove the `print(x)` line, both Python 2 and 3 don't hoist the enclosing `x` into `locals()` (since it can't see a single usage of that variable), resulting in a `KeyError`: Traceback (most recent call last):
File "blah.py", line 7, in <module>
outer()
File "blah.py", line 5, in outer
inner()
File "blah.py", line 4, in inner
print("x = {x}".format(**locals()))
KeyError: 'x'
Python 3 has the `nonlocal` keyword that you can use to indicate that you're using a variable from an enclosing scope so that it's properly introduced into `locals()` but Python 2 doesn't have this facility.Non-funky behaviour would be to have an `enclosing()` that walks up the outer functions and returns the union of their `locals()`. Alternatively, a `vars()` which expands to all variables in scope (respecting LEGB) would be best given the context of variable interpolation.
Also, if this thing is truly mangling bytecode, it's not portable between
different python versions [...] seems to be done by interpy_untokenize,
which boils down to some string mangling
It uses the Python file encoding property ("# coding: foobar") to rewrite the source code, not the bytecode, and they refer to pyxl as an inspiration.For a good explanation, see https://github.com/dropbox/pyxl
I'm not the poster but IMHO it fixes Python's syntax which is a case of DRY violation.
EDIT: To me, seeing this is Python code anywhere would seem to violate principle of least astonishment, which I think is somewhat more important than being DRY, if that's even a problem here.
"Hello {person}, it's a {weather} day".format(person=person, weather=weather)
The list of variables (person, weather) shows up three times.Somewhat off-topic, a similar problem shows up when you have a bunch of related functions, all taking a particular kwarg (or kwargs), and calling each other. Like
def foo(arg, conf1=None, conf2=None):
bar(arg+2, conf1=conf1, conf2=conf2) # eww :(
def bar(arg, conf1=None, conf2=None):
# etc.
A neat thing that perl6 has is syntax for "keyword argument whose value is found in the variable of the same name". So the equivalent of that line could be written bar($arg+2, :conf1($conf1), :conf2($conf2))
but it could also be bar($arg+2, :$conf1, :$conf2) cache.add_message("Hey, your {zoop} is {boop}".format(zoop=zoop, boop=boop))
That that call to "format" just kind of sucks -- it repeats what's already pretty obvious by looking at the string. It's specially crappy when you have strings that need lots of variables. We've taken to doing this recently, which I'm usually okay with: cache.add_message("Hey, your {} is {}".format(zoop, boop))
The downside of that is you have to make sure that the order of the arguments matches exactly with the order of the empty brackets. It's kinda error prone... But generally not that big of a deal.We could also do this...
cache.add_message("Hey, your %s is %s" % zoop, boop))
Or this... cache.add_message("{1} alert! Hey, your {0} is {1}".format(zoop, boop))
So yeah, I would argue that string formatting in Python is ALREADY in a kinda nasty place. There's ALREADY a bunch of ways to do it, and it all just kinda sucks.IMO, Ruby-style string formatting is probably the nicest I've seen. If it were in Python, it absolutely would be THE way to do string formatting, I bet.
cache.add_message("Hey, your {zoop} is {boop}")
So much nicer. bar='batz'
"foo #{bar} #{0.5+0.5}"
=> "foo batz 1.0"
"foo %s %.2f" % [bar, 1.0]
=> "foo batz 1.00"So I'm happy to see this, even if I'm probably too conservative to use it in my day job. And I didn't know about the coding: thing, and it looks like this method could also be used on my other python-wtf, which makes me even happier.
(My other python-wtf is that there really ought to be nicer syntax for a['b']. For a while I thought that a::b would be nice, but then I remembered that that could be a slice, so it can't be parsed reliably. a$b is probably my next choice. Or even require that kind of slice to be written with a space or something, like "a: :b".)
[1] Requiring braces even for a simple variable name seems like a poor decision. There's a little-known language called Haxe which IIRC gets it right: you can embed variables with just "hello $foo", or expressions with "your score is ${kills-deaths}". I get that Ruby allows unusual characters in variable names, and it's not obvious whether "is this yours, #name?" means #{name?} or #{name}?. But I'd rather have that potential for confusion than force the braces even when there's no ambiguity.
irb(main):006:0> @foo = 'bar'
=> "bar"
irb(main):007:0> "This is the value of foo: #@foo"
=> "This is the value of foo: bar">>> name = "Foo Bar"
>>> age = 25
>>> "Hi, my name is {name} and I'm {age} years old.".format(splatlocals())
"Hi, my name is Foo Bar and I'm 25 years old."
Arguably, this an abuse of `locals()`, but it gets you very nearly the same kind of use-variables-in-strings-with-curly-braces functionality.
Edit: HN markdown doesn't seem to let you escape the star italics operator. To be clear, you have to double-star splat locals().
https://github.com/ekimekim/pylibs/blob/master/libs/interpol...
It's not quite as natural as your one:
def foo(x):
print interpolate("Hello, {x}")
Though I do actually prefer having the explicit formatting call there so I know when the interpolation is being performed. Side effects and all that. In a perfect world, this is the syntax I'd prefer: def foo(x):
print "Hello, {x}".format()
ie. a format() without args defaults to "all variables accessible in the current scope".
I wouldn't actually want it to support arbitrary python the way ruby does, I find the .format() syntax flexible enough.(Also, my current implementation is for locals only. It wouldn't be hard to extend to globals, but would suffer the "nonlocals won't be captured" problem described in other comments here no matter what)
(EDIT: Also, it relies on sys._getframe, which is CPython specific)
I really enjoyed Ruby String interpolation, and "".format(...) or "" % (...) seems very verbose to me. I'm lazy by nature ;) "My Adam is, I am 10 old."
Correcting your example: "My name is {name}, I am {years} years old".format(name=name, years=years)
So to throw that out, that one line includes the word "name" four freaking times, and years four freaking times. You say you like the verbosity of it. Why? Would you like this format yet more? "My name is {name=name}, I am {years=years} years old".format(name=name, years=years)
If not why not? It's yet more verbose.I think that the reason that other people like the non-verbose format of:
"My name is {name}, I am {years} years old" # assuming the presence of local variables "name" and "years"
Is that, well, it's pretty obvious what's going on here, and repeating name and years a bunch more times do not, it seems, make it any more clear what's going on.A reasonable argument might be that:
"My name is {name}, I am {years} years old".format()
Is more clear about what's going on. But repeating the variable names is not particularly elucidating.https://www.python.org/dev/peps/pep-3101/
I guess finding discussion of the % formatting would be harder. I suspect that the discussion would have been about the tradeoff between the implicit variable insertion and simpler positional examples:
"My name is %s, I am %s years old" % (name, years)
(where there is definitely at least a tendency to avoid implicit behavior in the design of python; of course positional formatting like that is implicit, but it is quite a bit less implicit than automatically pulling variables out of the current scope)edit: the % formatting was probably informed by sprintf.
from timeit import Timer
try:
from StringIO import StringIO
except ImportError:
from io import StringIO
nr = 1200000
data = "The Quick Brown Fox Jumps Over The Lazy Dog: Woven silk pyjamas exchanged for blue quartz.\n"
# contruct a list first, then join
def dolist():
s = []
a = s.append
i = 0
while i < nr:
a(data)
i += 1
s = "".join(s)
print("%s chars (joined list)" % len(s))
# string concatenation fest
def dostr():
s = ""
i = 0
while i < nr:
s += data
i += 1
print("%s chars (string concatenation)" % len(s))
# use a string as a file
def dostringio():
buf = StringIO()
w = buf.write
i = 0
while i < nr:
w(data)
i += 1
s = buf.getvalue()
print("%s chars (cStringIO)" % len(s))
if 1:
tlist = Timer("dolist()", "from __main__ import dolist")
print("the joined list took %.2f seconds" % tlist.timeit(2))
tstr = Timer("dostr()", "from __main__ import dostr")
print("the concatenation fest took %.2f seconds" % tstr.timeit(2))
tlist = Timer("dostringio()", "from __main__ import dostringio")
print("the cStringIO approach took %.2f seconds" % tlist.timeit(2))
else:
@profile
def callall():
# For use with a profiler (eg, kernprof.py/lineprof)
for i in xrange(2):
dolist()
dostr()
dostringio()
callall()
Result:
(user@air) /Users/user/Prj/python $ python3 stringplakbenchmark.py
109200000 chars (joined list)
109200000 chars (joined list)
the joined list took 1.12 seconds
109200000 chars (string concatenation)
109200000 chars (string concatenation)
the concatenation fest took 1.76 seconds
109200000 chars (cStringIO)
109200000 chars (cStringIO)
the cStringIO approach took 1.45 seconds
(user@air) /Users/user/Prj/python $ python2.7 stringplakbenchmark.py
109200000 chars (joined list)
109200000 chars (joined list)
the joined list took 0.99 seconds
109200000 chars (string concatenation)
109200000 chars (string concatenation)
the concatenation fest took 1.33 seconds
109200000 chars (cStringIO)
109200000 chars (cStringIO)
the cStringIO approach took 5.21 seconds
(user@air) /Users/user/Prj/python $ python2.6 stringplakbenchmark.py
109200000 chars (joined list)
109200000 chars (joined list)
the joined list took 0.95 seconds
109200000 chars (string concatenation)
109200000 chars (string concatenation)
the concatenation fest took 1.39 seconds
109200000 chars (cStringIO)
109200000 chars (cStringIO)
the cStringIO approach took 5.54 secondshttps://github.com/dropbox/pyxl
Reminds me of JSX/E4X/XML Islands, etc.
Just a doubt: how do I specify source file encoding if the coding string is now hijacked for interpolation purposes? Is there a default, fixed encoding (which I hope is not iso-8859-1, python's own default)?
For interpy, I think it assumes utf-8 by default.
https://github.com/syrusakbary/interpy/blob/master/interpy/c...
Think of possibilities like
# coding: JIT
or
# coding: inline-C
etc.
https://docs.python.org/3.5/library/ast.html#ast.NodeTransfo...
the coding way you can invent any wild syntax.