Yes, the built in csv module really does that in Python 3.
If you pass it bytes, yes, it does. If you pass it strings, no, it doesn't.
If what you pass to the built-in CSV writer is not a string, the CSV writer will call str() to get a string representation it can write out. The string representation of a bytes object includes the 'b' prefix.
Meanwhile, you discovered your bug: you were treating bytes as text, which is likely to blow up on you sooner or later, and thanks to how Python now handles text, it blew up on you immediately as a way to remind you not to treat bytes as text.
What you probably think you want is for the CSV writer to realize it got a bytes object and, instead of calling str(), call its decode() method to get text it can write. But that is once again a dangerous operation, and sort of the whole point of Python 3's text changes is it won't let you get away with that stuff anymore.
It doesn't "blow up". If it blew up and retired, I would have seen the problem. The problem, like several python 2/3 incompatibilities, is that Python 3 merrily did something different, without telling anyone, until eventually we track down what has changed. I spent quite a while on this very bug myself, and it, along with others, persuaded me to switch to a different language (serious I know, but I was just getting annoyed with python's general loose dynamic nature, in combination with the python 2/3 changes.)
It's not a bug, it's a fix for an architectural error in Python 2, and it was quite well announced at the time: https://docs.python.org/3.0/whatsnew/3.0.html
But the fact that the same code now silently does the always wrong thing in Py3 wrt CSV is clearly a bug.
Actually, the design defect here is calling str() on everything, and assuming that the output is sensible for CSV. It may be a decent rule of thumb, but it clearly does not apply to bytes. Given the likelihood that someone might mistakenly use bytes as a string (for example, because they're porting a legacy Py2 codebase), this should be a hard error, immediately reported as such, and not just a silent behavior change.
Or do we say "CSV outputs strings, whatever is the string representation of what you passed in is what gets written out", and trust people to figure out when they're working with something that has a "wrong" string representation for their use case?
Because remember: the whole underlying cause of this was treating a dangerously non-string value as a string. Those bytes objects should have been decoded to strings long before reaching the CSV writer. Python 3 does raise more and louder exceptions when you pass bytes to things that expect strings, but the CSV writer isn't a thing that expects strings; it expects things that have a string representation, and several common use cases get much more difficult if you change that to force every user to explicitly do throwaway casts to string in the name of protecting people who keep insisting on writing dangerous "I'll treat bytes as string until it breaks, and then complain that the language did the wrong thing, not me" code.
So yeah, I would be fine with making an exception for it (and providing some kind of option to disable that exception, for that incredibly rare case where someone really does need b"foo" in their CSV output).
And then in 5 years, flip the default of that switch, and deprecate it. In another 5, remove it entirely.
Also, note that raising an error in this case is not placating the people who insist on using bytes as strings. Quite the opposite - it very loudly and unambiguously tells them that they're wrong, and how exactly they're wrong.
Um, this is exactly what I proposed above!
"this should be a hard error, immediately reported as such, and not just a silent behavior change."
What I'm asking for is that csv writer raises an exception if it sees bytes anywhere by default. The problem is that right now, it doesn't! It just gives you "incorrect" output, that might go undetected for a long time.
I think the Python maintainers fundamentally disagree with you on that point, but a fix submission would settle it unambiguously.
Sure, don't change the stuff that works and cause yourself unnecessary pain, but don't blame python 3 for your misencoded data.
If a company tries Python 3 and discovers basic things like CSV produce utter gibberish, they would do well to opt out. And they do--in droves.
My data is not misencoded, you see. It's just misunderstood.
So I won't be able to use NumPy with either of my two native languages. That sounds like a bit of a shortcoming for the majority of the world.
It could be argued that the csv module's behaviour is reasonable, and NumPy's isn't. (I'm not 100% sure about all the details of this issue) Hopefully, NumPy will change it's behaviour to match Python 3, but if not you could still use the NumPy CSV routines like `loadtxt` or `genfromtxt` [0]. So then this becomes a documentation change to add some warnings to both modules.
> they would do well to opt out. And they do--in droves.
This is simply not true. They would do well to handle strings properly and so avoid bugs in future - something Python 3 actively encourages, and Python 2 obscures. And while I can't speak for every company, our metrics show that our Python 3 code has far less customer issues than Python 2, Perl, or Ruby. Now that's business value. (Edit: I mean it's hard to make the comparison - the Perl code is e.g. older - but we're writing code now, and when the interns add new stuff to the Python 3 codebase, it breaks less. All of them are still actively developed, and the Ruby one is about as old as the Python 3 one).
[0] https://docs.scipy.org/doc/numpy/reference/generated/numpy.l...
Text encoding issues are absolute garbage in Python 3.x
I fucking hate the way that csv module works with text encodings.
As soon as I can figure out a reliable way to take latin-1 and save it as UTF-8 without breaking everything, I will try to shoehorn in a PR.
Right now, it's fucking awful. My ETL pipeline hates it, I hate it, my boss hates it, and my internal constituents hate it. Because it sucks.
A file I can read in one encoding and write as another should be readable with the encoding I wrote it in. That is not currently the case with the latest version of Python.
And it makes me hate the world.
with open('some latin-1 file', 'rb) as f:
text = f.read().decode('latin-1')
with open('some utf8 file', 'wb') as f:
f.write(text.encode('utf-8'))
Python 3's string encoding support is super good. I've said it before and I'll say it again: if you use bytes as a string you are Doing It Wrong.If you use bytes as a string you are Doing It Wrong.
If you use bytes as a string you are Doing It Wrong.
I do that operation on a file I get from an API. I know for a fact that the encoding I'm receiving is latin-1.
I run exactly that operation on the file that you wrote out in code.
When I try to read that file back in as UTF-8, I get encoding errors. That does not make for "super good." That makes me want to scream.
I do not have this problem when I use Python 2.7.x
To be perfectly clear: bytes (b'') is not a string. Again: bytes is NOT a string. It is an array of octets, aka bytes, aka unsigned 8 bit integers. NOT characters. NOT a string.
If you are dealing with bytes that are encoded representations of a string, then you have to know what encoding they use to decode them and treat them as strings.
in_file = open("in.csv", 'r', encoding='latin-1')
in_csv = csv.reader(in_file)
out_file = open("out.csv", 'w', encoding='utf-8')
out_csv = csv.writer(out_file)
out_csv.writerows(row for row in in_csv)The resulting file is not readable, and it makes me want to kick puppies and punch kittens.
I don't see how silently printing a binary literal, if that is indeed what it does, is reasonable. Simply put, b"foo" is not meaningful CSV.
What it should do is 1) raise an exception by default, informing the user that they need to be supplying strings and not bytes, and 2) provide an explicit switch to treat binary data as pass-thru, which would be useful in scenarios where you're just reading a file and dumping it elsewhere, and don't want to spend time decoding and then encoding everything.
+ if (PyBytes_Check(field)) {
+ append_ok = FALSE;
+ Py_DECREF(field);
+ PyErr_SetString(PyExc_TypeError, "Field is bytes");
+ }
else {
This would then raise a TypeError.I don't think this is the right solution. It seems weird to have a special case because people aren't watching what they're putting in. Garbage in, garbage out, consenting adults and all that.
But "practicality beats purity".
And "errors should never pass silently"!
It's the same kind of thing that leads to safety warning stickers put on products. You may read it and think that it's something so obvious that consenting adults should know better. But then you look at the statistics about how many people did not, and realize that, yeah, a sticker along the lines of "don't stick your finger into a food processor" is actually a good idea. Especially given how cheap it is, and how expensive reattaching fingers is...
Basically, products should be designed around known human weaknesses, and that includes entrenched modes of thinking by past products. It doesn't mean that new products should accommodate those entrenched modes, especially when they lead to other problems. But they should try to detect them, and issue clear and explicit warnings, to guide the person to the proper way of doing things.
In fact, the reason I chose Python was because I was able to dive so quickly into real problems like this with no problems whatsoever.
This is a minor issue for someone porting from 2 to 3, it is not a problem with 3
Python's a superlative language but it has a pretty terrible set of included libraries. urllib2 isn't the only library with a superior alternative on pypi. Pretty much all of them do.
I hate Python 3's removal of the (lambda (key, value): blah) tuple unpacking syntax, and the forcing of parentheses for print statements. They might seem minor but they aren't for me. So I'm not at all eager to move to version 3 and don't really see any benefit. Not sure if those who aren't migrating feel the same way, but I wouldn't be surprised if some of them do.
(Edit to address comment below: There are more issues I have with Python 3. It allows more bugs to slip through, for instance. I actually particularly like a comment I just wrote, so I'll link to it here: https://news.ycombinator.com/item?id=13145299 Do note that this was added after the reply below.)
I meant I don't see any benefits for me, not benefits for other people. I assumed that was clear; sorry if it wasn't.
> Unicode.
Yeah, but some people have still been living without the changes, and it's hardly enough of a reason on its own (for me anyway) when there's other things I hate about the language.
> Async.
It's a nice feature, yeah. I can live without it, as people have for many years. Maybe if I was used to having it around I wouldn't want to go back, but I'm not.
> Extended library.
Cool! I'm not sure what exactly falls under this that I'm supposed to be missing, but pip install has sure been taking care of everything in the blink of an eye in version 2.
> Required keywords.
Cool! I need it about as much as I need a donut.
> syntax inprovements (lots of m, many more than just the removal of the print statement).
nonlocal is literally the only positive one I can think of right now that I'd actually care about. But then again, it comes up maybe 50x less often than the parentheses I have to write for print, or the tuple unpacking that I have to do. So yeah, it's hardly a reason to migrate.
> Type hinting.
Nice to have. I'm living just fine without it. Maybe I'd have migrated if it actually optimized things or did something more useful.
You should read some changelogs of past python 3 releases. 3.6, for example, has ordered dicts by default. Which is quite convenient when you need to write test testing a small dict with two items for example.
I like driving an old muscle car, most of m look beautiful and bring me everywhere i want. But damn, those new cars changed a lot and are much more comfortable. (But they do break as much ;))
https://mail.python.org/pipermail/python-dev/2016-September/...
But it is insertion order indeed, not sorting order.
I think you just don't comprehend the multitude of problems that Python 3 fixes by handling strings correctly... Maybe you've never handled Unicode before.
First of all, the tuple unpacking one is a HUGE readability AND maintainability issue; it's not just syntactic sugar. var[0][1][2] is not only far less readable than the unpacking notation, it doesn't even have the same semantics (doesn't enforce the structure of the tuple).
That means Python 2.7 helped me catch more bugs. Think about that!
Second, no, I just listed the two that irritated me the most every time I tried to switch, because they were the first things that came up by far the earliest and most frequently. Small inconveniences can be amplified through their frequencies. There are lots of things I don't like about it though... division becoming floating point division, having to say list(foo.items()) or list(map(...)) instead of just foo.items(), etc... again, more verbosity and typing for common cases where I really didn't mind the old way. If I wanted imap(), I could've just used imap; they could've just moved that to __builtin__ and made my life easier that way.
By the way -- the lazy nature of map(), etc. also means you catch fewer bugs now. Again, think about that! Just because it looks more efficient, that doesn't mean it's actually better. If there's anything I've learned, it's that even the smallest things tend to come with non-obvious tradeoffs.
Finally, regarding strings: if you look at my earlier comments, yes, I already acknowledged the Unicode changes were for the better. Awesome. I agree. Cool? OK, but there are other things in the language besides Unicode though, and they're not as awesome. I don't spend my entire programming life dealing with Unicode strings, so I care about other things too, and they make my life harder. Simple as that.
I'm a huge lover of Python, but I drop into C# when I need stuff like that. Or Go. Or Rust. Or something.
You're just making bad decisions here.
If you're looking to catch bugs in your code before you test or deploy it, don't use a dynamic language.
Python has never been and probably never will be a language with declarative types and the checking that allows.
Pick a different hammer if that's the nail you need to hit. Don't complain about the hammer you want to use not being a screwdriver.
def foo((birthname,surname)):
...
you can write in Python 3: def foo(name):
birthname, surname = name
...
It's not less readable. I also missed it in the beginning, but it's not really a big deal. (I think they couldn't keep the feature because of how '*' is used, but I am not sure.)Edit: If you have problem with this in lambda expression, just create a named inner function. It's a feature/shortcoming (depends on POV) of Python that you cannot bind variables in an expression. I hope you understand that they couldn't keep the feature in lambdas if they didn't keep it in proper functions.
I am sure if you think about other things, there are good reasons to do it the way Python 3 does it, usually there is a hidden case where things need to be disambiguated (like your list() examples).
In any case, I think I see your problem. You are not the sole user of Python language. There are features that other people like (such as using '*' in unpacking), and so features you like are weighted against their use cases, and a reasonable compromise is made.
And frankly, I think if you like to use lambdas that much, you really want to program in a language where everything is an expression, such as Lisp or Haskell.
I'm glad I'm not. Otherwise I probably wouldn't be using it either. Not sure how that is my "problem".
> There are features that other people like (such as using [asterisk] in unpacking), and so features you like are weighted against their use cases, and a reasonable compromise is made.
I like that [asterisk] syntax too.
> And frankly, I think if you like to use lambdas that much, you really want to program in a language where everything is an expression, such as Lisp or Haskell.
Or I could just keep using Python 2.7 which works just fine, and not move to version 3 where I'm not welcome.
It's your problem in e.g. where you want list returned by default where Python 3 returns an iterator by default. Why is that useful for many people was already explained.
> I like that [asterisk] syntax too.
Funny, AFAIK it is Python 3 only.. https://www.python.org/dev/peps/pep-3132/
> Or I could just keep using Python 2.7 which works just fine, and not move to version 3 where I'm not welcome.
You are welcome to use Python 3, but - suit yourself. :-)
Misreading comments is not helping.
I think it's unfair to say that I misread his comment - he doesn't explicitly mention he is aware of the workaround I outlined for the functions, and that he is bothered with lack of tuple unpacking in lambda expressions only, not in ordinary functions.
Regardless, I still think it's quite impolite to downvote somebody who wants to help you and misunderstands you, if they are not e.g. factually incorrect. If you don't actually tell me where I am wrong, I cannot improve my answer. Also, this is not Stackoverflow, where that could be marginally acceptable (I am very strongly against downvoting without explanation).
Well, I don't consider it a "misunderstanding" when there are literally just 2 things to note in my comment that you're replying to ("lambda" and "tuple unpacking") and you still somehow miss 1 of them. I think it totally deserves a downvote, because it makes me look stupid when you present a reasonable solution to a non-problem and make readers assume I was saying something other than I was, and on top of that I have to waste some 5-10 minutes of my time replying. That's not something I appreciate.
That said, like I said, I never actually downvoted that comment (because I obviously couldn't). So you don't need to worry about the internet points.
I admit I don't use lambdas that much, since generator expressions (which is like Python 2.3) they aren't really needed too frequently. And in most cases you're better off using function anyway, because in Python statements are not expressions, as I already also stated. For example, I use print() for debugging frequently and this is tough to insert into lambda. (And even in Haskell I prefer to name subexpressions to lambda syntax.)
> That said, like I said, I never actually downvoted that comment (because I obviously couldn't). So you don't need to worry about the internet points.
I am not worried about internet points (I actually got about 80 of them on this discussion alone, which is frankly ridiculously too much, and in practice, I find that comments I personally find to be the most insightful only rarely get most points), I am just really annoyed when somebody downvotes my comments without any explanation, because I am a very curious person and in most cases it's just a honest misunderstanding, which could be cleared up with, I don't know, actual communication?
And at least two or three other people actually downvoted my original comment, so I would like to use this opportunity to invite them to come forward with an explanation what they found so wrong about it.
list(map(lambda (some, thing): some + thing, everything))
# better
list(some + thing for (some, thing) in everything)
Or, as the parent suggests, just create helper function, preferably one your python environment doesn't need to set up every time your outer function is called # okish, "verbose lambda"
def compute(everything):
def magic(elem):
some, thing = elem
return some + thing
return list(map(magic, everything))
# probably better
def magic(some, thing):
return some + thing
def compute(everything):
return list(magic(some, thing) for (some, thing) in everything)Too much nonsense in your comment. Really now? How about you give a realistic example where syntactic sugar doesn't substitute for it? Like what am I supposed to pass to sort(key)? And incidentally, this tuple unpacking problem comes up when sorting frequently... and that's the prime example on their web page (and a realistic one at that) for where you're supposed to use lambdas: https://docs.python.org/3/tutorial/controlflow.html#lambda-e... If you're telling me this isn't Pythonic, you're really just not being sensible.
I suggest you look at namedtuple: https://docs.python.org/3/library/collections.html#collectio...
I think you should watch some Raymond Hettinger's talks, he is discussing many little things like this.
This has been the general trend in mainstream languages lately, not just in Python. E.g. in C#, all LINQ operations are lazy. in Java, the new stream API, to be used with lambdas, is lazy.
In cases where the APIs were there, they were generally not as easily accessible (i.e. they were the equivalent of imap etc, with some hoops to jump before you could use them). The new APIs are more straightforward to use.
The reason to change the behavior of an existing API is because the default (i.e. most obvious) API should also be the most flexible, and do the right thing in as many cases as possible. This was not the case with map etc in Py2.
The disadvantage of changing an existing API like that is that it breaks code. But Py3 broke code anyway, so it was a good time to introduce breaks like that for the sake of better defaults.
Array.FindAll, Array.Convert, etc. all existed in C# beforehand. Though maybe this is what you meant in the next sentence.
> In cases where the APIs were there, they were generally not as easily accessible (i.e. they were the equivalent of imap etc, with some hoops to jump before you could use them).
This is going on a tangent but LINQ still has hoops to jump through. You have to say "using System.Linq;" at the top if you want to use the new syntax. That's like saying "from itertools import *" and then using imap, which you could've always done.
Yes, it's what I had in mind. I have to admit that I completely forgot about ConvertAll (and so assumed there was no map).
I think the biggest reason why those weren't all that commonly used in practice, is because in .NET you often deal with opaque collection types (like ICollection<T>, or ReadOnlyCollection<T>, or even custom-made collections pre-generics) that are usually exposed on properties of objects. Since the concrete type is not known, you can't do List.ConvertAll etc.
This, by the way, is another point in the favor of lazy implementations - they don't care about input type, because the output type is always "lazy sequence". Of course, you can have an eager map similarly not care about input, but then what should be the type of its output collection by default? No matter what type you choose, someone will complain that they wanted someone else. Given Python's preference for explicitness, such design would warrant several functions like map_to_list, map_to_tuple, map_to_set etc. But, of course, if you have a lazy map, you might as well just write list(map(...)) etc.
> You have to say "using System.Linq;" at the top if you want to use the new syntax. That's like saying "from itertools import " and then using imap, which you could've always done.
It's a bit different, though. When you import itertools, it brings all those functions into your global namespace. But when you import System.Linq, it only brings one static class into your global namespace; the actual functions are extension methods that only show up on the types to which they are applicable. So the resulting namespace pollution is far less in C#.
There's also the issue of import being generally frowned upon in idiomatic Python, largely because the way conflict resolution works there (silent override). In C#, if you happen to have clashing identifiers from usings, it'll prevent you from using them unqualified, so there's no good reason to avoid it.
You mean print functions. I love the new change because you can pass "print" around like any other function now, letting you write code like:
def my_map(data, func):
for item in data:
func(item)
my_map(dataset, insert_into_database)
# For testing
my_map(dataset, print)If your software is being actively maintained, it's time to move to Python 3.
Maintenance is key. Most people don't stick around for 20 years anymore either. I know I'm going to have an easier time finding a new hire for a Python codebase. And he's going to have a far better chance at understanding said codebase. Code which nobody knows how to maintain will hurt us either with a fiendish bug, or limit out growth. So for me, slowly moving away from legacy stuff is good business value in the long run.
Remember, you can never be sure that Fortran code is 100% bug free. The test of time is as good as any other test, but not perfect.
(FWIW, I'd probably write most new numerical code in Julia rather than Fortran 20xx, and either call into existing Fortran via FFI, or drive it from the command line with some scripting language.)
SciPy has Fortran code under the hood: https://github.com/scipy/scipy/search?l=FORTRAN&utf8=%E2%9C%...
Both SciPy and NumPy use LAPACK - a Fortran library. SciPy also uses BLAS - another Fortran library.
So every time you praise Python for being useful for scientific work, you're actually praising Fortran and C libraries/modules wrapped in Python.