I'm not sure what sane behavior Python could have here besides errorring.
> EVERY language should _try_ to handle Unicode such that if a data sequence were valid before it remains valid after.
This sequence was never valid, and never will be.
> in the article's case, the correct answer is GIGO. Just pass it through and hope it continues to work.
Dear God, no; emit a diagnostic and abort. Countless decades of existing code have shown time and again that "plow forward with some hot garbage" is not a good idea. But that ignores that … that that isn't how any of this works; the YAML parse is going to want to emit strings, which the incoming data isn't.
Neither were four-byte UTF-8 characters at some point.
> and never will be.
We shall see.
Does it work if you set the environment variable PYTHONENCODING to cp1252?
(I suppose I should either contact the author, or try it myself...)
If you Google, "Mikołaj",
> Mikołaj is the Polish cognate of given name Nicholas
Then Google, "windows character encoding polish"
> Windows-1250 - Wikipedia
And 0xb3 is "ł" in that encoding.¹
> Does it work if you set the environment variable PYTHONENCODING (sic) to cp1252?
I don't know if setting PYTHONIOENCODING would work here; I don't think it should affect this. Really, fixing the YAML file is the fix. (And fixing the thing that generated it.)
¹it is queries like this that really make me love the search engines of today. This would have been hell in the days of Alta Vista.
And this is why you always validate your data when you slurp it in, or else you pass crap down several layers where it crashes or mostly works with the potential for security holes or catastrophic behavior, and a pain in the arse to track down since the actual bug is nowhere near where you are looking.