Migrating to Python 3 with pleasure
github.com
github.com
In particular, being able to create an updated copy of a dict with a single expression is pretty cool:
return {**old, 'foo': 'bar'}
# Old way
new = old.copy
new['foo'] = ['bar']
return new [*a, *b, *c] In [1]: x = {1:2, 3:4}
In [2]: %timeit x[3] = 5
48.6 ns ± 1.18 ns per loop (mean ± std. dev. of 7 runs, 10000000 loops each)
In [3]: %timeit y = x.copy(); y[3] = 5
189 ns ± 3.23 ns per loop (mean ± std. dev. of 7 runs, 10000000 loops each)
In [4]: %timeit {**x, 3: 5}
182 ns ± 3.3 ns per loop (mean ± std. dev. of 7 runs, 10000000 loops each)
Edit: It also seems to be pretty constant time if you're just merging: In [16]: %timeit {**x, **y, **z}
180 ns ± 1.29 ns per loop (mean ± std. dev. of 7 runs, 10000000 loops each)
In [17]: %timeit {**x, **y, **z, 3: 5}
278 ns ± 18.2 ns per loop (mean ± std. dev. of 7 runs, 1000000 loops each)
In [19]: dis.dis(lambda: {**x, **y, **z, 3: 5})
0 LOAD_GLOBAL 0 (x)
2 LOAD_GLOBAL 1 (y)
4 LOAD_GLOBAL 2 (z)
6 LOAD_CONST 1 (3)
8 LOAD_CONST 2 (5)
10 BUILD_MAP 1
12 BUILD_MAP_UNPACK 4
14 RETURN_VALUE
In [20]: dis.dis(lambda: {**x, **y, **z})
0 LOAD_GLOBAL 0 (x)
2 LOAD_GLOBAL 1 (y)
4 LOAD_GLOBAL 2 (z)
6 BUILD_MAP_UNPACK 3
8 RETURN_VALUE return {**old, 'foo': 'bar'}
# Old way
return dict(old, foo='bar')
Not much difference if you ask me. {**x, 'fo'+'o': 'bar', **y}Personally I hate the new {} for set syntax.
Is {} an empty set, or an empty dict?
X ++ {'fo'+'o': 'bar'} ++ y
Where '++' here means dictionary union, the choice of symbol is not relevant.
return dict(x, **{'fo'+'o': 'bar'}, **y)
The new syntax is some mild syntactic sugar. (Which isn't a bad thing IMO)The expanded unpacking works in all cases.
Just skip all this and use Perl instead. You can write far more idiomatic, succinct readable code in Perl than you can in Python.
The whole point of Python is to not write code this way.
Python 2.7.12 (default, Dec 4 2017, 14:50:18)
[GCC 5.4.0 20160609] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> x = {'foo': 5, 'bar': 6}
>>> y = {'foo': 7, 'baz': 9}
>>> dict(x, **{'fo'+'o': 'bar'}, **y)
File "<stdin>", line 1
dict(x, **{'fo'+'o': 'bar'}, **y)
^
SyntaxError: invalid syntax
Python 3.5.2 (default, Nov 23 2017, 16:37:01)
[GCC 5.4.0 20160609] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> x = {'foo': 5, 'bar': 6}
>>> y = {'foo': 7, 'baz': 9}
>>> dict(x, **{'fo'+'o': 'bar'}, **y)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: type object got multiple values for keyword argument 'foo'
>>> {**x, 'fo'+'o': 'bar', **y}
{'foo': 7, 'baz': 9, 'bar': 6}dict(a, ... ,b) feels cleaner.
baz = {**foo, **bar}a, *b, c = range(10)
car, *cdr = some_list
(OMG I fired up the Python REPL and you totally can!)For anyone having trouble with maintaing multiple Python versions, I recommend Pyenv. You can install multiple local versions and switch between them. The selected version then uses the standard commands "python" and "pip", etc. which you can use to make your virtualenv from.
Also, your phrasing invites the question. Why on earth would I use a new feature in a turing complete language? Unless using the new feature results in tangible improvements in my code, I don't see a reason to use it.
In fact, there are lots of reasons not to use an extra language feature that doesn't provide tangible improvements in code, including maintainability, ease of reading code, and portability in python's case.
One surprising thing I learned from this document is that dicts now iterate in assignment order, not hash order. That's going to break some code for people.
Might be a bit contentious but it's not supposed to be relied upon
(Source: https://mail.python.org/pipermail/python-dev/2017-December/1...)
Among other things this makes writing tests much easier.
Starting in 2.7.3 and 3.3, the hashing algorithm applied a per-process random salt to certain built-in types, in order to mitigate a potential denial-of-service vector by crafting dictionary keys which cause hash collisions. Unless the PYTHONHASHSEED environment variable was set to a consistent value, iteration over a dictionary was no longer accidentally consistent across runs.
In 3.6, the dictionary implementation was rewritten to drastically improve performance. As a side effect, dictionary iteration became consistent again, this time determined by insertion order. In 3.6, this consistency was an implementation detail and not to be relied on.
Beginning with 3.7, dictionary iteration order is finally guaranteed by the language, independent of implementation, and goes by insertion order.
Raymond Hettinger, Modern Python Dictionaries... hopefully this is the correct one. https://youtu.be/npw4s1QTmPg
From a theoretical angle, the only way to build hash tables with some kind of provable runtime guarantees is to include randomness.
> Keys and values are iterated over in an arbitrary order which is non-random, varies across Python implementations, and depends on the dictionary’s history of insertions and deletions. If keys, values and items views are iterated over with no intervening modifications to the dictionary, the order of items will directly correspond.
That is, the order will be maintained between calls, so long as the dictionary had not been modified.
Going back to "str(some_dict) == str(some_dict)". I would not expect to always be the same, but for entirely different reasons. Consider:
class Strange():
def __repr__(self):
import random
return str(random.random())
some_dict = {1: Strange()}
>>> str(some_dict) == str(some_dict)
FalseI should have quoted the next section:
> If items(), keys(), values(), iteritems(), iterkeys(), and itervalues() are called with no intervening modifications to the dictionary, the lists will directly correspond.
Curiously, the "Dictionary view objects" section at https://docs.python.org/2.7/library/stdtypes.html#dictionary... has the same "Keys and values are iterated over in an arbitrary order which is non-random ..." text, but without being inside of a "CPython implementation detail" box.
I'd expect that any two ways of forming a dict with exactly same contents can result in different str(x) representation, and also that serializing and deserializing a dict can result in a different string representation.
It's mad that it ever wasn't this way. Mapping-with-ordered-keys is such a useful and pervasive data structure (all database query result rows, for one) that an ordered dictionary should be a fundamental part of a language.
It has been so much more pleasant to write python since ordering was maintained by default.
What? No. SQL does not return results in any consistent ordering unless specifically instructed to.
> Not the result set, the rows of the result set.
Each row is a mapping with ordered keys.
A result set has rows, which are not in a deterministic order unless an "order by" is provided. Each row has columns. The columns are in order, obviously.
I said that database query result rows are made up of ordered dictionaries. I didn't mention ordering of the result set.
Happy to keep repeating this as many times as necessary.
That said, you disagreed with my question then went on to show my question was on point.
The “rows themself” being an ordered map means you are referring to the columns, the order being set by the SELECT clause or table definition order (in case of wildcard).
That said, I personally feel iterating over table columns in that way to be a “bad code smell”. Not saying it’s bad in all cases, but generally it’s an anti-pattern to me.
"An array which represents an n-ary relation R has the following properties:
1. Each row represents an n-tuple of R.
2. The ordering of rows is immaterial.
3. All rows are distinct.
4. The ordering of columns is significant -- it corresponds to the ordering S1, S2, ..., Sn of the domains on which R is defined (see, however, remarks below on domain-ordered and domain-unordered relations).
5. The significance of each column is partially conveyed by labeling it with the name of the corresponding domain."
-- A relational model of data for large shared data banks[1]
[1]: https://cs.uwaterloo.ca/~david/cs848s14/codd-relational.pdf
I didn't mention the collection of rows, I mentioned the rows themselves.
> The “rows themself” being an ordered map means you are referring to the columns
No, it means I'm referring to the rows themselves.
The rows themselves are each ordered mappings.
Hope this helps to clarify things. Happy to keep repeating this as many times as necessary.
but hey, its Guido, I still can't fathom that he moved reduce into functools.
https://mail.python.org/pipermail/python-dev/2016-September/...
Here's one real-life use case of that: Avro records. One of the formats Avro uses is a text format that's basically ordered JSON. One company I worked for years ago used Avro as its wire protocol, and some Avro data was stored as JSON files on disk. Of course, Python's JSON implementation by default loads/unloads JSON to/from a dict. So just calling json.load() and json.dump() means I can't just load an Avro record from disk, change some data, and save it (which is something that came up when I was writing an upgrade script at a company I was working at years ago).
Thankfully, I had an out: the JSON library lets you override what container you load JSON into with object_pairs_hook, so I could just snarf it into an OrderedDict. But if I ever have to do this again after 3.7 comes out, I'm glad I won't have to worry about making sure I have the right container class. It makes my code simpler, and I won't have to leave a comment explaining why the code will break unless I specify an OrderedDict.
I hate to think of all the developer hours wasted because JSON doesn't maintain key ordering.
OrderedDict is fundamentally different because it is designed to allow inserting/removing keys in the middle of the ordering in O(1) time. It does not use the new dict under the hood, or vice-versa.
aside from that, python has a separate ordered dict class
Personally, I'm worried people will come to rely on the new behaviour in code instead. As core developers have repeatedly said, dict order is still an implementation detail, it should not be relied on as it is not officially part of the language. Other implementations (except pypy) will probably not have this behaviour.
Yet, I feel like this will fall on deaf ears and become a de-facto part of the language. Blogs will state it as a new feature, Python books will teach it and new coders will rely on it, forever locking the dict internals in place for all python interpreters.
(If you need to rely on ordering, use an OrderedDict instead.)
If you were relying on keys order before, you were not only doing it wrong, but you were doing so despite being told again and again.
This is of course pure speculation...
Current FreeBSD Python 3 status: https://wiki.freebsd.org/Python
I'm glad I did. F-strings are wonderful, as is pathlib and the enhanced unpacking syntax.
Since I started my current job, I've also been writing as many scripts as I can in Python 3 as well (and Docker has been a godsend for that because I can now deploy 3.6 scripts to the few servers we have that are still running RHEL 6).
I'll often do something similar to this, where I have a CLI tool that I'm not ready to deploy server-wide yet, and has hefty dependencies.
The pattern I use is to have a wrapper shell script that calls:
docker run -it --rm --volume "$(pwd)/$1:/file_to_process:z" --user $(id -u) container-image /opt/command_to_run /fileToProcess
This runs "/opt/command_to_run /fileToProcess" inside a container as the current uid, mounting the parameter to the shell script as "/file_to_process" inside the container.The :z mount parameter may or may not be needed depending upon whether you have SELinux enabled or not (by default, SELinux prevents countainers accessing any file on the host, and :z changes the SE context to allow access). I don't know if this is the case with EL6 tho.
The -t parameter shouldn't be used if your script is running in a pipeline (it creates a pseudo-tty). So it may be worth having some kind of conditional to remove this.
The wrapper I use also has a conditional to add the "$(pwd)" prefix to the call parameter only if the parameter is a relative path.
Then I push the image to my company's internal Docker registry, ssh into the server, and pull the image.
(also, aside from using Python 3 on RHEL 6, it also means using Python 3.6 on RHEL 7 without having to install python36u)
Edit:
Dockerfile (script name redacted)
FROM python:3
WORKDIR /usr/src/app
COPY requirements.txt ./
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
CMD [ "python", "./redacted.py" ]
Build and push commands (company and script names redacted): docker build --rm -t dreg.example.com/redactedproject/redacted .
docker push dreg.example.com/redactedproject/redacted
Pull and run on the server (same stuff redacted as above, plus I redacted the actual port number to be on the safe side): docker pull dreg.example.com/redactedproject/redacted
docker run -p 1337:1337 -d --name redacted dreg.example.com/redactedproject/redacted
And at some point, I'll make proper startup scripts for them. On RHEL 7 boxes, I've made systemd unit files. On RHEL 6... well, I suppose I'll be writing initscripts soon.I don't want to have to start teaching pyenv to every academic or student I want to send a script to.
It has been really useful for me in that I want the advantages of type annotations and f-strings and other python 3.5+ features but I have to support running in an environment with only 2.6 installed. So when I target a 3.5+ version, all of those features are maintained, but when I target 2.6, the transpiler does all the work in converting to running 2.6 code for me.
Thanks for the useful list!
I think many people get imprinted with writing the PY3 code using only the features available at the time they switched over.
And collections.ChainMap.
And f-strings.
And yield from.
And type hints.
And statistics.
And ipadress.
And secrets.
And matmul.
And subprocess.run.
Come on, Python 3 is packed with awesomeness !
What you have is decision making analogous to the function
evaluate(string_interpolation, python_ecosystem)
which is different from evaluate(string_interpolation)This argument won't fly. Did you know that string interpolation in Perl and shell was always an opt-in feature? And despite that it was frowned upon by Python community.
While certainly awesome, these are in python 2.7 as well. Although I don't remember if they where in 2.7 first of back-ported to 2.7 from 3.
> Previously it was always tempting to use string concatenation (concise, but obviously bad), now with pathlib the code is safe, concise, and readable.
This is the kind of feature that I'm wary to use even in scripts: questionable benefit, and probably too clever.
TypeError: expected str, bytes or os.PathLike object, not int >>> from pathlib import Path
>>> p=Path(1)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "/usr/lib/python3.6/pathlib.py", line 979, in __new__
self = cls._from_parts(args, init=False)
File "/usr/lib/python3.6/pathlib.py", line 654, in _from_parts
drv, root, parts = self._parse_args(args)
File "/usr/lib/python3.6/pathlib.py", line 638, in _parse_args
a = os.fspath(a)
TypeError: expected str, bytes or os.PathLike object, not int
>>>
So what happens if the paths are numbers? They are treated like any other characters.Python is strongly typed.
Should I have said "numeric characters" or "numeric strings" in my last paragraph instead of numbers?
>>> Path.cwd() / 1
TypeError: argument should be a path or str object, not <class 'int'>
Also, note that the '/' operator is left-associative, so even if you have two adjacent numbers when joining paths, there will be no division.It's not like programs consist of path operations to a significant degree, so using fancy syntactic sugar doesn't seem like a worthwhile optimization to me. At all.
I almost think it'd be nice to make matmul undefined for ranks higher than 2, since it's not really matrix multiplication and if you want to do that (or the previous behavior of dot) it can be achieved with einsum, with the advantage that you have to be a lot more explicit about what sort of tensor multiplication you want. That's probably a bit too purist though.
Homebrew, apt, pacman, etc. all have one-line python3 installation.
train_path = datasets_root / dataset / 'train'
[1] https://docs.python.org/3/library/pathlib.html
[2] https://docs.python.org/3/library/pathlib.html#basic-use
Except the boolean operators (and, or, not). For instance __and__ overrides the binary and (&), not the boolean and.
[1] https://docs.python.org/3/reference/datamodel.html#object.__...
The 'typing' module in the standard library was new as of 3.5.
Syntactic support for annotating variables was new as of 3.6.
Support for delaying resolution of annotations is new in 3.7 with a __future__ import.
Originally the annotation feature was seen as a possible way to add type hints to Python, but other potential uses were envisioned and no immediate preference was given to types over other uses of annotations.
Mypy supports it: http://mypy.readthedocs.io/en/latest/python2.html