`Bytes`: The Lesser-Known Python Built-In Sequence
thepythoncodingstack.com
thepythoncodingstack.com
I didn't think so, but then interviewing Juniors (who all invariably have "expert python" on their CVs), nobody ever seems to know them, or anything about string handling/unicode.
Are the basics about how stuff is moved around inside computers just completely opaque to people learning to program nowadays? I learned that stuff at secondary (high) school, so internally my brain just counts it as "foundational knowledge".
Really it's just a case of Python being able to handle a ton of different use cases that sometimes overlap, you can be really deep into it in one area but never touch another.
Even when these topics are covered directly in a university course, this knowledge doesn't stick around to become "foundational". Especially for those that don't have a hardware, embedded, or otherwise resource-constrained background or area of interest.
> who all invariably have “expert python” on their CVs
This is to get around HR filters, where they ask your experience with various tools on a 1-10 scale. If you’re only seeing the “experts” it’s possible the people who used more reasonable assessments of their abilities were simply rejected.
Early Python was ASCII-oriented. Type "str" was used both for text strings and arrays of bytes. "str" was treated much like arrays of "char" in C - it was the basic type for binary I/O.
Then came Unicode and Python 3. Characters and bytes were no longer the same thing. Strings and arrays of bytes had to be split somehow. As with most languages that had to retrofit Unicode, this didn't go well.
Rust, which didn't have a retrofit problem, did a clean separation. There are arrays of u8, and there is "str", which must be valid UTF-8 sequences. There is no implicit conversion. This is straightforward.
Python didn't do it that way. Type "bytes" in Python 3 prints as an ASCII string, not an array of numbers. It's close to the 'str" from Python 2. The usual string-like operations, such as "split" are defined for "bytes". This was kind of weird but was supposed to simplify conversion of old code from Python 2 to Python 3. It didn't really help all that much.
There's also "bytearray", which is a mutable version of "bytes".
It's one of those messes left over from the ASCII to Unicode transition era, along with UTF-16, byte order marks, "wchar_t", HTML character sets, HTTP headers, and all the character representations in SQL.
The whole decade-long mess was just to fix quirks in Windows' path and file encoding handling bullshit. And they didn't even manage to fix the problem in the end.
What a mess.
Totally agree. One example of "python-brain" that I see a lot at my employer is using pandas dataframes as the only data structure, beyond _maybe_ the occasional list or dict. Any entity with multiple attributes is represented as more columns in a dataframe, rather than via objects.
Chances are somewhat good that these people weren't computer science majors to begin with. For example, math or biology majors who have moved away from R to Python might know a great deal about data but not much about compsci.
For people who use Python in a DevOps context, they'll likely be exposed to more OOP concepts and lean more heavily on classes.
After 5 years of Go I might reasonably have been called an expert (I started using Go in 2012)—there wasn’t much left in the language to learn apart from the things that had changed. But with Python there’s tons and tons left for me to learn even ignoring all of the changes.
Just looking at this list of modules: https://docs.python.org/3/library/index.html
I reckon I've used at least once probably 50% of the list, but that's over 12 years of using Python, and I'd only consider myself fluent with a very small number of the total.