One surprising thing I learned from this document is that dicts now iterate in assignment order, not hash order. That's going to break some code for people.
One surprising thing I learned from this document is that dicts now iterate in assignment order, not hash order. That's going to break some code for people.
Starting in 2.7.3 and 3.3, the hashing algorithm applied a per-process random salt to certain built-in types, in order to mitigate a potential denial-of-service vector by crafting dictionary keys which cause hash collisions. Unless the PYTHONHASHSEED environment variable was set to a consistent value, iteration over a dictionary was no longer accidentally consistent across runs.
In 3.6, the dictionary implementation was rewritten to drastically improve performance. As a side effect, dictionary iteration became consistent again, this time determined by insertion order. In 3.6, this consistency was an implementation detail and not to be relied on.
Beginning with 3.7, dictionary iteration order is finally guaranteed by the language, independent of implementation, and goes by insertion order.
Might be a bit contentious but it's not supposed to be relied upon
(Source: https://mail.python.org/pipermail/python-dev/2017-December/1...)
Among other things this makes writing tests much easier.
Raymond Hettinger, Modern Python Dictionaries... hopefully this is the correct one. https://youtu.be/npw4s1QTmPg
Personally, I'm worried people will come to rely on the new behaviour in code instead. As core developers have repeatedly said, dict order is still an implementation detail, it should not be relied on as it is not officially part of the language. Other implementations (except pypy) will probably not have this behaviour.
Yet, I feel like this will fall on deaf ears and become a de-facto part of the language. Blogs will state it as a new feature, Python books will teach it and new coders will rely on it, forever locking the dict internals in place for all python interpreters.
(If you need to rely on ordering, use an OrderedDict instead.)
From a theoretical angle, the only way to build hash tables with some kind of provable runtime guarantees is to include randomness.
I'd expect that any two ways of forming a dict with exactly same contents can result in different str(x) representation, and also that serializing and deserializing a dict can result in a different string representation.
> Keys and values are iterated over in an arbitrary order which is non-random, varies across Python implementations, and depends on the dictionary’s history of insertions and deletions. If keys, values and items views are iterated over with no intervening modifications to the dictionary, the order of items will directly correspond.
That is, the order will be maintained between calls, so long as the dictionary had not been modified.
Going back to "str(some_dict) == str(some_dict)". I would not expect to always be the same, but for entirely different reasons. Consider:
class Strange():
def __repr__(self):
import random
return str(random.random())
some_dict = {1: Strange()}
>>> str(some_dict) == str(some_dict)
FalseI should have quoted the next section:
> If items(), keys(), values(), iteritems(), iterkeys(), and itervalues() are called with no intervening modifications to the dictionary, the lists will directly correspond.
Curiously, the "Dictionary view objects" section at https://docs.python.org/2.7/library/stdtypes.html#dictionary... has the same "Keys and values are iterated over in an arbitrary order which is non-random ..." text, but without being inside of a "CPython implementation detail" box.
It's mad that it ever wasn't this way. Mapping-with-ordered-keys is such a useful and pervasive data structure (all database query result rows, for one) that an ordered dictionary should be a fundamental part of a language.
It has been so much more pleasant to write python since ordering was maintained by default.
Here's one real-life use case of that: Avro records. One of the formats Avro uses is a text format that's basically ordered JSON. One company I worked for years ago used Avro as its wire protocol, and some Avro data was stored as JSON files on disk. Of course, Python's JSON implementation by default loads/unloads JSON to/from a dict. So just calling json.load() and json.dump() means I can't just load an Avro record from disk, change some data, and save it (which is something that came up when I was writing an upgrade script at a company I was working at years ago).
Thankfully, I had an out: the JSON library lets you override what container you load JSON into with object_pairs_hook, so I could just snarf it into an OrderedDict. But if I ever have to do this again after 3.7 comes out, I'm glad I won't have to worry about making sure I have the right container class. It makes my code simpler, and I won't have to leave a comment explaining why the code will break unless I specify an OrderedDict.
I hate to think of all the developer hours wasted because JSON doesn't maintain key ordering.
What? No. SQL does not return results in any consistent ordering unless specifically instructed to.
That said, you disagreed with my question then went on to show my question was on point.
The “rows themself” being an ordered map means you are referring to the columns, the order being set by the SELECT clause or table definition order (in case of wildcard).
That said, I personally feel iterating over table columns in that way to be a “bad code smell”. Not saying it’s bad in all cases, but generally it’s an anti-pattern to me.
"An array which represents an n-ary relation R has the following properties:
1. Each row represents an n-tuple of R.
2. The ordering of rows is immaterial.
3. All rows are distinct.
4. The ordering of columns is significant -- it corresponds to the ordering S1, S2, ..., Sn of the domains on which R is defined (see, however, remarks below on domain-ordered and domain-unordered relations).
5. The significance of each column is partially conveyed by labeling it with the name of the corresponding domain."
-- A relational model of data for large shared data banks[1]
[1]: https://cs.uwaterloo.ca/~david/cs848s14/codd-relational.pdf
I didn't mention the collection of rows, I mentioned the rows themselves.
> The “rows themself” being an ordered map means you are referring to the columns
No, it means I'm referring to the rows themselves.
The rows themselves are each ordered mappings.
Hope this helps to clarify things. Happy to keep repeating this as many times as necessary.
> Not the result set, the rows of the result set.
Each row is a mapping with ordered keys.
A result set has rows, which are not in a deterministic order unless an "order by" is provided. Each row has columns. The columns are in order, obviously.
I said that database query result rows are made up of ordered dictionaries. I didn't mention ordering of the result set.
Happy to keep repeating this as many times as necessary.
but hey, its Guido, I still can't fathom that he moved reduce into functools.
https://mail.python.org/pipermail/python-dev/2016-September/...
OrderedDict is fundamentally different because it is designed to allow inserting/removing keys in the middle of the ordering in O(1) time. It does not use the new dict under the hood, or vice-versa.
aside from that, python has a separate ordered dict class
If you were relying on keys order before, you were not only doing it wrong, but you were doing so despite being told again and again.