Advanced Python Mastery
github.com
github.com
David wrote https://www.dabeaz.com/generators/ which remains one of my all-time favourite Python tutorials. Looking forward to digging into this.
It's 8 years old, but definitely worth a watch: https://www.youtube.com/watch?v=MCs5OvhV9S4
"A fantastic, entertaining and highly educational talk. It always bothers me that I can't play the piano and talk at the same time (my wife usually asks me things while I'm playing). But David can even type concurrent Python code in Emacs in Allegro vivace speed and talk about it at the same time. An expert in concurrency in every sense of the word. How enviable!"
Moreover, this video is from PyCon. Which is sponsored by PSF. Which is sponsored by google, meta, AWS etc. So who is pocketing the ad money? Why do I want to pay to support that person? As well as Google and its advertisers?
The logic in your comment is incoherent.
My favorite introductory book (not an introduction to programming but an introduction to the language) is “Introducing Python by Lubanovic” because it’s one of the only beginner books that actually covers the python module system with enough depth and the second half of the book gives a quick overview of a lot of different python libraries.
1. Test-Driven Development with Python
2. Architecture Patterns with Python
The 2nd one is the closest you're gonna get to a production-grade tutorial book.
Related to this topic, these resources by @dbeazley:
Barely an Interface
https://github.com/dabeaz/blog/blob/main/2021/barely-interfa...
Now You Have Three Problems
https://github.com/dabeaz/blog/blob/main/2023/three-problems...
A Different Refactoring
https://github.com/dabeaz/blog/blob/main/2023/different-refa...
His youtube channel:
-> Currently serving as Application Architect for a medium sized Python application.
[0] https://github.com/dabeaz/blog/blob/main/2023/three-problems...
inp is just standard in python, along with uin (user input), or sometimes also raw or iraw.
You cannot win this battle. There is nothing wrong with 'inp'
> Have you actually done code reviews professionally?
Yes, I did.
In a serious code review, that isn't even a starter of an issue, if you have context surrounding that variable. Furthermore `input` is an actual Python function, and shortening an example for learning purposes is not the same as asking other people do the same in production code.
> lambda has the benefit of making the code compact and foreboding. Plus, it prevents people from trying to add meaningful names, documentation or type-hints to the thing that is about to unfold.
Disclaimer, I did not read the entire post.
It's 47 minutes and totally worth it.
The course materials for this course and the introductory course (“Practical Python”[1]) are quite thorough, but I've always found the portfolio analysis example very hokey.
There's enormous, accessible depth to these kinds of P&L reporting examples, but the course evolves this example in a much less interesting direction. Additionally, while the conceptual and theoretical materials is solid, the analytical and technical approach that the portfolio example takes quickly diverges from how we would actually solve a problem like this. (These days, attendees are very likely to have already been exposed to tools like pandas!) This requires additional instructor guidance to bridge the gap, to reconcile the pure Python and “PyData” approaches. (Of course, no other Python materials or Python instruction properly address and reconcile these two universes, and most Python materials that cover the “PyData” universe—especially those about pandas—are rife with foundational conceptual errors.)
Overall, David is an exceptional instructor, and his explanations and his written materials are top notch. He is one of the most thoughtful, most intelligent, and most engaging instructors I have ever worked with.
I understand from David that he rarely teaches this course or Practical Python to corporate audience, instead preferring to teach courses direct to the public. (In fact, I took over a few of his active corporate clients when he transitioned away from this work, which is what led me to drafting my own curricula.) I'm not sure if he still teaches this course at all anymore.
However, I would strongly encourage folks to look into his new courses, which cover a much broader set of topics (and are not Python-specific)! [2]
Also, if you do happen to be a Python programmer, be sure to check out his most recent book,“Python Distilled”[3]!
[1] https://dabeaz-course.github.io/practical-python/
[2] https://www.dabeaz.com/courses.html
[3] https://www.amazon.com/Python-Essential-Reference-Developers...
Well, unfortunately 5-day courses listed there are $1500 each.
If the free of charge course discussed here is really that good, it is a nice promo to go and pay for another. Ed Tech Lo-Fi style.
This is an example from the linked course https://github.com/dabeaz-course/python-mastery/blob/main/Ex...:
``` # readport.py
import csv
# A function that reads a file into a list of dicts
def read_portfolio(filename):
portfolio = []
with open(filename) as f:
rows = csv.reader(f)
headers = next(rows)
for row in rows:
record = {
'name' : row[0],
'shares' : int(row[1]),
'price' : float(row[2])
}
portfolio.append(record)
return portfolio
``` def read_portfolio(filename):
with open(filename) as f:
rows = csv.reader(f)
headers = next(rows)
for row in rows:
yield {
'name' : row[0],
'shares' : int(row[1]),
'price' : float(row[2]),
}
Now you can call read_portfolio() to get an iterable that lazily reads the file and yields dicts: portfolio = read_portfolio()
for record in portfolio:
print '{shares} shares of {name} at ${price}'.format_map(record) header, *records = [row.strip().split(',') for row in open(filename).readlines()]
but then you need a way to parse the records, which could be Template() from the string library or something like... type_record = lambda r : (r[0], int(r[1]), float(r[2]))
At this point, the two no longer mesh well, unless you would be able to unpack into a function/generator/lambda rather than into a variable. (I don't know but my naive attempts and quick SO search were unfruitful.) Also, you're potentially giving up benefits of the CSV reader. Plus, as others have clarified, brevity does not equal readability or relative lack of bugs:In the course example, it's reasonably easy to add some try blocks/error handling/default values while assigning records, giving you the chance to salvage valid rows without affecting speed or readability. In fact, error handling would be a necessity if that CSV file is externally accessible. Contrast that with my two lines, where there's not an elegant way to handle a bad row or escaped comma or missing file or virtually any other surprise.
Anything else I can think of off-hand (defaultdict, UserList, "if not portfolio:") has the same initialization step, endures some performance degradation, is more fragile, and/or is needlessly unreadable, like this lump of coal:
portfolio = [record] if 'portfolio' not in globals() else portfolio + [record]
So... your technique and generators. Those are safe-ish, readable, relatively concise, etc.> header, *records = [row.strip().split(',') for row in open(filename).readlines()]
Better would be:
header, *records = [row.strip().split(',') for row in open(filename)]
No need to read the lines all into memory first.Edit: Also if you want to be explicit with the file closing, you could do something like:
with open(filename) as infile:
header, *records = [row.strip().split(',') for row in infile]
That is if we wanted to protect against future changes to semantics for garbage collection/reference counting. I always do this, but I kind of doubt it will ever really matter in any code I write.It looks like that code does read the whole file:
(with a foo.csv that is 350955 bytes long:)
% python -V
Python 3.11.4
% python
>>> f = open("foo.csv")
>>> f.tell()
0
>>> header, *records = [row.strip().split(',') for row in f]
>>> f.tell()
350955
I thought that using a list comprehension to bind header and records was eagerly consuming the file, so I changed it to a generator comprehension with >>> f.close()
>>> f.open("foo.csv")
>>> header, *records = (row.strip().split(',') for row in f)
>>> f.tell()
350955
nope, I guess the destructuring bind does it? >>> f.close()
>>> f.open("foo.csv")
>>> headers, records = f.readline().strip().split(','), (row.strip().split(',') for row in f)
>>> f.tell()
125
not as neat, though. Is there a golf-ier way to do it?* [... for row in open(filename).readlines()]
The readlines return value is one copy, and the list comprehension is another copy. However, that first copy can be avoided with: [... for row in open(filename)]
The entire file must still be read to evaluate the list comprehension.Additionally, this doesn't do what you think it does:
>>> header, *records = (row.strip().split(',') for row in f)
Compare to this, using a variable for clarity: >>> gen = (row.strip().split(',') for row in f)
>>> header, *records = next(gen) def read_portfolio(filename):
record = lambda r: {
'name': r[0],
'shares': int(r[1]),
'price': float(r[2]),
}
with open(filename) as f:
rows = csv.reader(f)
headers = next(rows)
return [record(r) for r in rows]That will read the file as needed (ie as you iterate over it) instead of loading the entire thing in memory.
for record in read_portfolio(fn):
# do stuff def read_portfolio(filename):
with open(filename) as f:
rows = csv.reader(f)
headers = next(rows)
return [
{
"name": r[0],
"shares": int(r[1]),
"price": float(r[2]),
}
for r in rows
]You could use a list comprehension, but that can be unclear and hard to extend, depending on the situation. It can be a nice option if most of the parts in the generator can be broken out into functions with their own name, though.
You could turn it into a generator, which can cause some fun bugs (e.g. everything works fine when you first iterate over it, but not afterwards), so IMO that's best used when it needs to be a generator, for semantics or performance.
You could turn it into a generator, then add a wrapper that turns it into a list (keeping the inner function private), or use a decorator that does the same, but it's less clear than this pattern.
So, i'd just learn to live with it.
Yeah, so many people don't get this, but too many small functions can be hard to understand -- that's why I qualified that option.
In this case i agree that inlining it is fine, i was talking about the general pattern.
So `portfolio = [{'name': row[0], 'shares': int(row[1]), 'price': float(row[2]) for row in rows]`
But if it's more complicated than this (like if there is conditional(s) inside the loop), I'd recommend just stick with the current approach. It's possible to have even multiple conditionals in list comprehension, but it's not really very readable. If you do want to, walrus operator can make things better
(something like `numbers = [m[1] for s in array if (m := re.search(r'^.*(\d+).*$', s))]`)
keys = "name", "shares", "price"
portfolio = [
dict(zip(keys, row))
for row in rows
]
If I had to do more complex stuff than building a dict like this I'd move it into a function. That tends to make the purpose more clear anyway.That said, it's fine to append to a list too, I just prefer comprehensions when they fit the job. In particular, if you're just going to iterate once over this list anyway, you can turn it into a iterator comprehension by replacing [] by () and save some memory.
def get_record(row):
return {
'name': row[0],
'shares': int(row[1]),
'price': float(row[2])
}
return [ get_record(r) for r in rows ]
or return list(map(get_record, rows))In 2023, it's rude to return dictionaries. :)
Why do you say that?
I think you were partially kidding, but also half serious. What’s the issue with returning dictionaries, and why should we be returning dataclasses instead?
Asking for my own learning.
def some_method(input: Dict[str, int]):
...
You can just as easily call `some_method({"foo": 1})` as `some_method({"bar": 2})`.vs.
def some_method(input: MyDataClassWithFoo):
...
Now you can't pass a dict where the key is "bar" and presumably get a KeyError when it tries to look up a "foo" key, you can only pass a MyDataClassWithFoo.Regarding rude, yea it was a bit tongue-in-cheek (I would never hold a grudge against somebody for returning a dict).
You should generally define an interface for your function that's as precise as its logic allows for. `Dict` is as good as `Any`: it doesn't tell you very much about what the function internally expects. The sibling comment to mine does a good job going through this.
``` There should be one-- and preferably only one --obvious way to do it. ```
with open(filename) as f:
rows = csv.reader(f)
next(rows)
return [
{
'name': row[0],
'shares': int(row[1]),
'price': float(row[2]),
} for row in rows
]Personally the way you've done it is the most Pythonic IMO. List comprehensions are great but would be less readable in this case.
def iter_portfolio(rows):
for row in rows:
yield {'name': row[0]}
rows = ...
portfolio = list(iter_portfolio(rows))Honestly, this is the approach I've been using even though I hate it. Specially if your code is going to be read by anyone other than you.
However, because of available ML/DL/LLM frameworks and libraries in Python, Python has been my go to language for years now. BTW, I love the other comment here that Beazley is the Jimi Hendrix of Python. Only those of us who enjoyed hearing Hendrix live really can get this.
Please use ctypes, cffi or https://github.com/wjakob/nanobind
Beazley himself is amazed that it (Swig) is still in use.
I googled, and this seems to be the current edition:
https://www.amazon.in/Python-Essential-Reference-Essentia-De...
> total_cost = 0.0
> with open('../../Data/portfolio.dat', 'r') as f: > for line in f: > fields = line.split() > nshares = int(fields[1]) > price = float(fields[2]) > total_cost = total_cost + nshares * price
> print(total_cost)
yikes what a terrible reference implementation! Least they could do is reduce([...], +) as a two-liner
map/reduce (for better or worse) get a bad rap. Some blog or training I took when first starting up with the language told me 'map/reduce Bad' and I have generally avoided ever since.
David is also doing online immersive courses - 1 week long but I believe he also splits them into one day sessions now. Highly recommended !
He has a real talent for explaining complicated concepts in a very simple and approachable way.
But then again, I don't use list comprehensions, because I don't comprehend them, so what do I know.
Are the "people from diverse backgrounds" (whatever that means) denied any of their rights and privileges by a programmer's use of "advanced" code? No. Quite the contrary: less-experienced people may still read, learn from and improve themselves by the advanced/beyond-their-level code they encounter. So what's your real issue?
Code is written to solve problems; not to please the lowest common denominator. If you can't read code, you're that denominator, and that's on you.
I find your statement terribly odd too. A true Python master tends to write highly readable code; idiomatic and Pythonic code tends to be readable, unlike in other languages. So if you can't read advanced Python (or even use list comprehensions, which are basic in the Python scheme of things), you're not qualified to opine on what "advanced mastery" of Python entails.
Where did you get that idea? Python is a programming language. That some find it more accessible than others is orthogonal to the work needed to master it.
Software also has an ongoing quality crisis, all while being more and more influential in people's lives. That person's attitude towards writing quality software helps to deepen that crisis and is therefore harmful.
Not everyone is trying to be nice either.
To put it bluntly. I'd hire the downvoted guy/gal in a heartbeat but would shy away from the downvoters. Why? Because I need to deliver business value which pays for our salaries. And I need to do it today and tomorrow and in years from now.
This is a message to normal people that understand that coding is a _social_ activity that has an audience in the present (your coworkers) and in the future (poor maintainers). Not only you're not alone in this but you are the majority.
I appreciate the aesthetic beauty of great code. But it has a cost compared to average code.
This is doubly true for a language like Python, which occupies a niche of "lingua franca between users with wildly different backgrounds bringing value to the table by being able to use and change the same software".
So, to me, a Python zen master would not write incomprehensible code, but instead write readable code very quickly that effectively and efficiently solved the problem they are facing due to their comfort working inside the Python ecosystem.
(short but exact comments, inline testing, good function and variable naming, overall good use but not overuse of the standard library, functions very rarely more than a dozen lines, generally understandable code)
Why?
I reserve downvotes for posts that are flagrant, factually wrong, or are otherwise against the rules. Flagging might also be used. But using downvotes to have a voice not be heard feels wrong, too. What was said doesn't hurt anyone, even if a vast majority of people around here might disagree with it. Downvoting because you disagree feels wrong.
I disagree with this. But I can't downvote this comment because it is a reply to my comment. This restriction specifically exists because downvoting to disagree is how HN has always worked.
> highly upvoted comments don't get bolded the way that downvoted comments get grayed.
Highly upvoted comments float to the top and therefore have more visibility.
I don't think so? For instance, I can upvote replies to my comments, so it's not they have set up some "you can either comment or vote" system. This restriction is just a nudge toward positivity rather than negativity.
> Highly upvoted comments float to the top and therefore have more visibility.
Yes, but downvoted comments float down and are grayed out. The point is just that the two things aren't totally symmetrical.
(Also I think some of this stuff was implemented years after pg's pronouncement about downvotes.)
And listen, I didn't say "people who downvoted that comment aren't using HN correctly and should be booted off the site!". I just said "I think it is stupid to downvote that comment". And I do. It's stupid to downvote perfectly reasonable comments that you simply disagree with. Again, most people don't use HN that way, irrespective of anything pg said in 2008, or we would see a lot more gray comments, and I would have gotten a lot more downvotes over the years on stuff I've said that people disagree with, instead of comments telling me why I'm wrong.
For what it's worth, I contend that - notwithstanding what pg and dang said many years ago, this is the revealed preference of most HN users, because it's quite rare to see a comment that is downvoted, just because lots of people disagree with it.
This does not match my experience at all.
Advanced != incomprehensible. Incomprehensible isn’t a feature of advanced either, you can be a novice and still write incomprehensible Python code.
On Reddit, not HN: https://news.ycombinator.com/item?id=16131314
I don't consider that appeal to authority canonical. My opinion is that downvotes should be used for bad comments, not for comments you disagree with. These aren't the same thing.
You can't downvote me because my comment was in response to yours. It's not an appeal to authority. It's a statement of fact as also seen in the implementation itself.
It is the definition of an appeal to authority to quote an authority figure as the final word on some debate. I've been here since before that first comment from pg about this, and no, I don't agree that it's a "fact" that this is "how HN has always worked", and no, there is nothing in the implementation that unambiguously makes downvotes be for disagreement.
But again, I'm not arguing for strict rules that work the way I prefer. I'm saying, this is a community I participate in, and I have my own opinions about how best to participate in it, which aren't necessarily aligned with the people who created the site 15 years ago. Other people are entitled to differing opinions about this, and can use the site how they prefer, but I still have my own opinions and will advocate for them.
It's perfectly sane to actively avoid trying to understand the your tools of your trade better?
Now I recognise that we should, unless there's a very good reason (not for style), keep our code stupid-simple.
And doing that is harder than making it clever.
Oh man, that is some job security!
Generators seem unpythonic.
And what a disaster that has been -- a bunch of people now consider object-oriented programming to be a sensible approach.
Python's organic growth and adoption is more of an OpenCola model.