Yes, the built in csv module really does that in Python 3.
Yes, the built in csv module really does that in Python 3.
If you pass it bytes, yes, it does. If you pass it strings, no, it doesn't.
If what you pass to the built-in CSV writer is not a string, the CSV writer will call str() to get a string representation it can write out. The string representation of a bytes object includes the 'b' prefix.
Meanwhile, you discovered your bug: you were treating bytes as text, which is likely to blow up on you sooner or later, and thanks to how Python now handles text, it blew up on you immediately as a way to remind you not to treat bytes as text.
What you probably think you want is for the CSV writer to realize it got a bytes object and, instead of calling str(), call its decode() method to get text it can write. But that is once again a dangerous operation, and sort of the whole point of Python 3's text changes is it won't let you get away with that stuff anymore.
It doesn't "blow up". If it blew up and retired, I would have seen the problem. The problem, like several python 2/3 incompatibilities, is that Python 3 merrily did something different, without telling anyone, until eventually we track down what has changed. I spent quite a while on this very bug myself, and it, along with others, persuaded me to switch to a different language (serious I know, but I was just getting annoyed with python's general loose dynamic nature, in combination with the python 2/3 changes.)
It's not a bug, it's a fix for an architectural error in Python 2, and it was quite well announced at the time: https://docs.python.org/3.0/whatsnew/3.0.html
But the fact that the same code now silently does the always wrong thing in Py3 wrt CSV is clearly a bug.
Actually, the design defect here is calling str() on everything, and assuming that the output is sensible for CSV. It may be a decent rule of thumb, but it clearly does not apply to bytes. Given the likelihood that someone might mistakenly use bytes as a string (for example, because they're porting a legacy Py2 codebase), this should be a hard error, immediately reported as such, and not just a silent behavior change.
Or do we say "CSV outputs strings, whatever is the string representation of what you passed in is what gets written out", and trust people to figure out when they're working with something that has a "wrong" string representation for their use case?
Because remember: the whole underlying cause of this was treating a dangerously non-string value as a string. Those bytes objects should have been decoded to strings long before reaching the CSV writer. Python 3 does raise more and louder exceptions when you pass bytes to things that expect strings, but the CSV writer isn't a thing that expects strings; it expects things that have a string representation, and several common use cases get much more difficult if you change that to force every user to explicitly do throwaway casts to string in the name of protecting people who keep insisting on writing dangerous "I'll treat bytes as string until it breaks, and then complain that the language did the wrong thing, not me" code.
So yeah, I would be fine with making an exception for it (and providing some kind of option to disable that exception, for that incredibly rare case where someone really does need b"foo" in their CSV output).
And then in 5 years, flip the default of that switch, and deprecate it. In another 5, remove it entirely.
Also, note that raising an error in this case is not placating the people who insist on using bytes as strings. Quite the opposite - it very loudly and unambiguously tells them that they're wrong, and how exactly they're wrong.
Um, this is exactly what I proposed above!
"this should be a hard error, immediately reported as such, and not just a silent behavior change."
What I'm asking for is that csv writer raises an exception if it sees bytes anywhere by default. The problem is that right now, it doesn't! It just gives you "incorrect" output, that might go undetected for a long time.
I think the Python maintainers fundamentally disagree with you on that point, but a fix submission would settle it unambiguously.
Sure, don't change the stuff that works and cause yourself unnecessary pain, but don't blame python 3 for your misencoded data.
If a company tries Python 3 and discovers basic things like CSV produce utter gibberish, they would do well to opt out. And they do--in droves.
My data is not misencoded, you see. It's just misunderstood.
It could be argued that the csv module's behaviour is reasonable, and NumPy's isn't. (I'm not 100% sure about all the details of this issue) Hopefully, NumPy will change it's behaviour to match Python 3, but if not you could still use the NumPy CSV routines like `loadtxt` or `genfromtxt` [0]. So then this becomes a documentation change to add some warnings to both modules.
> they would do well to opt out. And they do--in droves.
This is simply not true. They would do well to handle strings properly and so avoid bugs in future - something Python 3 actively encourages, and Python 2 obscures. And while I can't speak for every company, our metrics show that our Python 3 code has far less customer issues than Python 2, Perl, or Ruby. Now that's business value. (Edit: I mean it's hard to make the comparison - the Perl code is e.g. older - but we're writing code now, and when the interns add new stuff to the Python 3 codebase, it breaks less. All of them are still actively developed, and the Ruby one is about as old as the Python 3 one).
[0] https://docs.scipy.org/doc/numpy/reference/generated/numpy.l...
In fact, the reason I chose Python was because I was able to dive so quickly into real problems like this with no problems whatsoever.
This is a minor issue for someone porting from 2 to 3, it is not a problem with 3
Text encoding issues are absolute garbage in Python 3.x
I fucking hate the way that csv module works with text encodings.
As soon as I can figure out a reliable way to take latin-1 and save it as UTF-8 without breaking everything, I will try to shoehorn in a PR.
Right now, it's fucking awful. My ETL pipeline hates it, I hate it, my boss hates it, and my internal constituents hate it. Because it sucks.
A file I can read in one encoding and write as another should be readable with the encoding I wrote it in. That is not currently the case with the latest version of Python.
And it makes me hate the world.
with open('some latin-1 file', 'rb) as f:
text = f.read().decode('latin-1')
with open('some utf8 file', 'wb') as f:
f.write(text.encode('utf-8'))
Python 3's string encoding support is super good. I've said it before and I'll say it again: if you use bytes as a string you are Doing It Wrong.If you use bytes as a string you are Doing It Wrong.
If you use bytes as a string you are Doing It Wrong.
To be perfectly clear: bytes (b'') is not a string. Again: bytes is NOT a string. It is an array of octets, aka bytes, aka unsigned 8 bit integers. NOT characters. NOT a string.
If you are dealing with bytes that are encoded representations of a string, then you have to know what encoding they use to decode them and treat them as strings.
I do that operation on a file I get from an API. I know for a fact that the encoding I'm receiving is latin-1.
I run exactly that operation on the file that you wrote out in code.
When I try to read that file back in as UTF-8, I get encoding errors. That does not make for "super good." That makes me want to scream.
I do not have this problem when I use Python 2.7.x
in_file = open("in.csv", 'r', encoding='latin-1')
in_csv = csv.reader(in_file)
out_file = open("out.csv", 'w', encoding='utf-8')
out_csv = csv.writer(out_file)
out_csv.writerows(row for row in in_csv)The resulting file is not readable, and it makes me want to kick puppies and punch kittens.
I don't see how silently printing a binary literal, if that is indeed what it does, is reasonable. Simply put, b"foo" is not meaningful CSV.
What it should do is 1) raise an exception by default, informing the user that they need to be supplying strings and not bytes, and 2) provide an explicit switch to treat binary data as pass-thru, which would be useful in scenarios where you're just reading a file and dumping it elsewhere, and don't want to spend time decoding and then encoding everything.
+ if (PyBytes_Check(field)) {
+ append_ok = FALSE;
+ Py_DECREF(field);
+ PyErr_SetString(PyExc_TypeError, "Field is bytes");
+ }
else {
This would then raise a TypeError.I don't think this is the right solution. It seems weird to have a special case because people aren't watching what they're putting in. Garbage in, garbage out, consenting adults and all that.
But "practicality beats purity".
And "errors should never pass silently"!
It's the same kind of thing that leads to safety warning stickers put on products. You may read it and think that it's something so obvious that consenting adults should know better. But then you look at the statistics about how many people did not, and realize that, yeah, a sticker along the lines of "don't stick your finger into a food processor" is actually a good idea. Especially given how cheap it is, and how expensive reattaching fingers is...
Basically, products should be designed around known human weaknesses, and that includes entrenched modes of thinking by past products. It doesn't mean that new products should accommodate those entrenched modes, especially when they lead to other problems. But they should try to detect them, and issue clear and explicit warnings, to guide the person to the proper way of doing things.
So I won't be able to use NumPy with either of my two native languages. That sounds like a bit of a shortcoming for the majority of the world.
Python's a superlative language but it has a pretty terrible set of included libraries. urllib2 isn't the only library with a superior alternative on pypi. Pretty much all of them do.