Tqdm (Python)
tqdm.github.io
tqdm.github.io
With the better error messages in 3.10 and 11, plus the focus on speed, it's a fantastic era for the language and it's a ton of fun.
I didn't expect to find back the feeling of "too-good-to-be-true" I had when starting with 2.4, and yet.
In fact, being a dev is kinda awesome right now, no matter where you look. JS and PHP are getting more ergonomics, you get a load of new low level languages to sharpen your hardware, Java is modern now, C# runs on Unix, the Rust and Go communities are busy with shipping fantastic tools (ripgrep, fdfind, docker, cue, etc), windows has a decent terminal and WSL, my mother is actually using Linux and Apple came up with the M1. IDEs and browsers are incredible, and they do eat a lot of space, but I have a 32 Go RAM + 1TO SSD laptop that's as slim as a sheet of paper.
Not to mention there is a lot of money to be made.
I know it's trendy right now to say everything is bad in IT, but I disagree, overall, we have it SO good.
Although the first release on pypi was in 2013: https://pypi.org/project/tqdm/#history
TIL tqdm now has a dead simple telegram hook https://tqdm.github.io/docs/contrib.telegram/
Please just admit you made a mistake instead of responding to everyone trying to defend your use of the word "new" on an old tool. :D
[0]: https://pydoit.org
I mean, it's declarative, works on Windows, easy things are (very) easy, and because you can mix and match bash and python actions, hard things are suspiciously easy too.
Given how complicated the alternatives are (maeven, ninja, make, gulp...), you'd think it would have taken the world for years.
Yet I've only started to see people in the Python core dev team use it this year. It's only getting traction now.
Care to share an example?
It has to make sure a bunch of directories exist, run a node js command to build some files from the proper dir, then run a python command to regroup them + all static files from all django apps into one dir. Simple.
But then I had a problem I had to hack around. This required me to change an entry to a generated TOML file on the fly at every build.
doit just lets me add a 5 lines python function that does whatever I want, and insert it between my bash tasks, and I'm done.
def task_bundle_static_files():
def update_manifest():
conf = Path("var/static/manifest.toml")
data = toml.loads(conf.read_text())
data["./src/main.js"] = data["index.html"]
conf.write_text(toml.dumps(data))
return {
"actions": [
"mkdir -p ./var/static/",
"rm -fr ./var/static/* ",
"cd frontend/; npm run build",
"python manage.py collectstatic --noinput",
update_manifest,
"cp -r ./var/static/* ",
]
}better yet, use both and just doit
btw, the project is old and it is useful with the current feature set. There is an issue with the bus factor but is common for many tools.
But I prefer it because:
- it runs anywhere you have python.
- it uses a clean syntax, and the parser is great at telling you about errors since it's python.
- you get access to python's stdlib: string formatting, maths, etc.
- you get access to python's ecosystem. Want to deal with timezone, hashing, crypto ? Sure you can.
- you get access to python's tooling (debugger, formatter, linter).
- you can still just use bash if you want. Easy things stay easy.
I can go home with Python at 5pm. And have a good time with my family.
Doing advent of code in Python is almost cheating. It feels like playing a video game.
My Python code in Part 2 (some minor adjustments from Part 1) took about 40 seconds to run. Terrible, but usable to get a proper answer. I was able to bring it down to ~25 seconds with some optimization by adding a calculated lookup dictionary per loop. Now, the same exact logic in Go ran in about 0.8 seconds.
However, I wasn't satisfied with this and realized that if I moved the dictionary to be global rather than per-loop, I would be able to realize significant gains in performance by completing eliminating redundant calculations. This change dropped the runtime from 25 seconds to 0.35 seconds (of course, applying the same logic in Go brought it down from 0.8s to 0.05s).
Due to the nature of performance you can get out of the Python interpreter, it can actually lead you down paths of learning better optimization strategies that you may initially write off in other languages (depending on the use case) because they perform inherently better. It made me think a bit more about what I was doing and how I could improve it since (in this particular case), the impact of not doing so was pretty drastic.
My Kotlin solution runs in about a second. And it was even so stupid that I didn't calculate the sum of the arithmetic series directly, but through a loop. Can't fathom something being slower.
line.split(",").map { it.toInt() }.let { crabs ->
(0..crabs.maxOf { it }).minOf { pos ->
crabs.sumOf {
(0..abs(it - pos)).sum() // slower than calculating arithmetic sum, but quicker to write
}
}
}
My proper Kotlin solution runs in less than a ms, though.I do think that "it's good that python is slow because it forces you to optimize" is a weird take, though.
My take wasn't that it forces you to, but that non-optimal code paths can be greatly exaggerated in comparison to other languages, particularly compiled ones. You can still ignore it (I mean, within reason), but it can give you that extra push to really look a bit deeper to understand what's going on. And of course, there's optimized libraries written in C/C++ that you can take advantage of for even better number crunching than standard CPython.
> What kind of dictionary, and how come even a naive solution would be so slow?
My naive solution was literally going through every single element for every loop and not storing any data besides the fuel buildup and the alignment number that generated it. The dictionary was added to act as a cache to store already computed fuel consumption values, initially per-loop then moved one level up to be global (because the summations wouldn't be different).
I'm not saying my method (posted in a sibling comment) is the best solution, but it's the way my brain walked through the problem.
Cool of you to participate without being a developer! Lots of computer science topics makes it easier, so hard without knowing of them. For instance graph searching / Dijkstra has beem relevant this week.
I'm not unfamiliar with programming, but I come from the sysadmin side of things. "Glue" work is usually where things are focused and the 'fun' nitty-gritty of algorithms can be a bit out-of-scope, though I'm not a sysadmin in my current role anymore so any dev-related work I do is purely personal now.
I've had to take a break from AoC, only got up through Day 10, but didn't get P2 for 8 and 9. It's a fun way to keep the mind going and to slip back into the coding space to at least not lose skills, even if the solutions are simple/non-optimal.
When I read the Day 7, I saw bruteforcing would lead to bad perfs. Then I remembered that in high school, I learned a formula to calculate the nth term of a sequence without having to process the entire sequence.
I couldn't remember the formula, nor the name of the concept, so I google around until I found some tutorials, and relearned what I was taught as a child: arithmetic sums.
The consumption for a crab can then be calculated in constant time:
def crab_consumption(crab, target):
n = abs(target - crab) - 1
return (n**2 + 3*n + 2) / 2
And the Python solution finishes instantly.Bottom line, I could keep using Python for all problems and benefit from the amazing productivity of it.
I've been using python since 2.4 now, and it's not always fast enough. But it very, very often, is.
def calculate_fuel(positions):
fuel = None
low_fuel_value = None
calculated = dict()
for value in range(positions[0], positions[-1]+1):
consumed = 0
for position in positions:
diff = abs(value - position)
try:
consumed += calculated[diff]
except KeyError:
if position != value:
consumption = sum([x for x in range(1,diff+1)])
calculated[diff] = consumption
consumed += consumption
if fuel:
if consumed < fuel:
fuel = consumed
low_fuel_value = value
else:
fuel = consumed
low_fuel_value = value
print(f"Aligning to: {low_fuel_value}")
return fuel
And Day 6 I fell for the bruteforce bait and had the thought "It can't be as easy as changing 80 to 256, right?". Then I realized the pain I had created for myself. BUT! My 6p2 code ran faster than my 6p1 by a good margin, which I was happy about. val crabs = lines.first().split(",").map { it.toInt() }
val avg = crabs.sum() / crabs.size
return crabs.sumOf { abs(it - avg).let {dst -> (dst * (dst + 1)) / 2} }
And similarly for part1 take the median. Why it kinda works:Part1: The median I felt made sense intuitively, as in my head I thought about an example ala 1,1,3,100. Never makes sense to use x>3, because even though the crab at x=100 then can walk shorter, there are 3 others then having to walk longer. And x=1,2or3 doesn't matter, just symmetrically changes which side has to walk one step less or one step more.
And for part2 I thought similar, except the cost is exponential and therefore I want to minimize the avg move and not the total moves, thus taking the average.
Doubly so this year, with the theme being linear algebra.
import numpy as cheat
:)I can never remember all the differences and subtleties between virtualenv, pipenv, venv, pyenv-virtualenv, workon, conda, and so on when I encounter them in a random git repo.
Jokes aside it's a great language. I would love to see a better package management ecosystem for it, as that is also my biggest issue. Python also does something no other mainstream scripted language does- it allows you to install extensions to the language right in the same package manager, as it can compile libraries from source when wheels are unavailable- this makes it a much harder challenge. At the same time I'm really happy that PyPI is a non-profit organization and won't have to go through the issues that something like NPM did.
- tqdm: progress bars (https://tqdm.github.io/)
- rich: text formatting (https://github.com/willmcgugan/rich)
- textual: TUI, using rich (https://github.com/willmcgugan/textual)
- fastpi: Rest APIs (https://fastapi.tiangolo.com/)
- typer: CLI Library, uses Click (https://typer.tiangolo.com/)
- pydantic: Custom data types (https://pydantic-docs.helpmanual.io/)
- shiv: Create Python zipapps (https://shiv.readthedocs.io/en/latest/)
- toga: GUI Toolkit (https://toga.readthedocs.io/en/latest/)
- doit: Task runner (https://pydoit.org/)
- diskcache: A disk cache (https://github.com/grantjenks/python-diskcache/)
FWIW python has a built-in disk cache—- shelve ( https://docs.python.org/3/library/shelve.html )
Those that aren't should be there momentarily :)
Python is great for programming at the speed of inspiration.
tests are super helpful and wise in any sort of long term or safety-critical software, obviously
so, tests (and Python) are not inherently a win or a loss. they just give you different trade-offs
times to write testless Python and times to write test-heavy Go, Rust etc.
I therefore try to put each situation quickly into 1 of 3 buckets then move forward on that basis:
1. heck no
2. heck yes
3. either way. a gray area
I generally see a case 1 and 2 with confidence. therefore by a process of elimination, that also lets me deduce when its case 3. and in those cases you cant go wrong. :)
"Impossible to predict, the future is." - Yoda
Python performance problems are exaggerated for many use-cases (hot paths are not written in Python e.g., matching regexes happens in C code(re,regex), parsing xml too (elementtree, lxml), sqlite, numpy,scipy (C, Fortran), etc. Cython makes it trivial to drop to C level where necessary)
Like looking at that very first example, I have no clue what "len" means in that context. Is it implicitly checking that it's not an empty string? Then on the next line, how come `int` has `Use()` around it, but on the previous line `str` didn't? I guess that int is being used as a converter, line on the next line with str.lower, but the str was being used as a type check?
Not a fan of that API.
Rich is just amazing library
No more bash! No more subprocess! Write shell-like code with the convenience of Python!
Seriously, it's a great module. A bit of a learning curve, but then it feels natural.
...bit of a learning curve?
Actually, for me Jupyter (I use lab but I'm sure NB is amazing too) is the tool that boosts my productivity the most. (And omg vim mode.)
Also have to mention the fantastic Python Prompt Toolkit, which xonsh is based on - https://python-prompt-toolkit.readthedocs.io/en/master/
I mirror another commenter's excitement on this thread about the cool libraries we have at this age, and the tools that are made possible by them.
It makes it a lot easier to casually use a database (e.g. SQLite) to persist dictionaries without explicitly building a schema.
Great for CLI apps and ad hoc data crunching.
Yet, you can still make 200K writing it, I don't know the next time I'll create a command line application in Python, but I'll keep this little tool in mind. I hope Python eats the world.
I just wish you would have provided ipynb instead of md. It would be so much fun to just modify and run as I am learning. Thank you!
https://github.com/powerpak/tqdm-ruby
(shameless plug and an invitation for pull requests)
pv and tqdm would look even better if they'd be called implicitly (with an opt-out) since I always end up regretting not using pv when my command is taking too long. Too late.
For reference, it's a backronym for 'Command Line Interface Creation Kit': https://click.palletsprojects.com/en/8.0.x/
I do wish I could hook into it to test better, the only thing you can really do right now is to have it print stuff out and assert the output string. It's not really necessary to just build a CLI with click, but I want to build a library that integrates with it and testing the integration is a PITA.
I want to write a config-loading library for CLI apps like Golang's Viper lib, but for Python
In Click projects I usually use that rather than tqdm directly.
My only complaint is the smoothing parameter; by default it predicts the estimated time remaining based on the most recent updates so it can fluctuate wildly; smoothing=0 predicts based on the total runtime which makes more sense given law of large numbers.
Something like the NO_COLOR environement variable?
These progress bars are nice when you launched a single loop yourself, but when you are running an automated battery of many things they become annoying and pollute your terminal too much. I know that you can silence each particular program that uses tqdm by setting a 'disable' option. But this requires editing the python source code of the program.
import os
import sys
from tqdm import tqdm as _tqdm
def tqdm(*args, **kwargs):
try:
disable = bool(int(os.environ['NO_PROGRESSBARS']))
except KeyError:
disable = not sys.stdout.isatty()
except (ValueError, TypeError):
disable = False
kwargs.setdefault('disable', disable)
return _tqdm(*args, **kwargs)
Then import that instead of tqdm.tqdm.sys.stdout.isatty() isn't a perfect answer to what people ask when they want to know "am I running in an automated environment, or is a human user looking at my output?", but it's close. More nuance is available online.
> import that instead of tqdm.tqdm
But it's not me who is importing tqdm to begin with! I call many programs in parallel from shell scripts (out of my control) and they all call tqdm individually. I need to stop tqdm output from outside these programs.
Your code should be part of tqdm itself, not written by individual programmers.
> am I running in an automated environment, or is a human user looking at my output?
But I want to stop tqdm output precisely because I'm a human looking at it. If you have more than one or two progress bars simultaneously, it becomes useless clutter.
As much as I like to use tqdm myself for my programs, I'm sad that as tqdm becomes more and more popular, my terminal output becomes more and more cluttered, to an absurd amount. Piping the output to a file does not help and is totally the wrong idea. I'm precisely interested in seeing--in real time--the part of the output that does not come from tqdm, such as warnings and errors.
My Python skills have improved so much thanks to HN.
Another great alternative is fastprogress: https://github.com/fastai/fastprogress
It often works better in Jupyter Notebooks.
It makes the module so much funner to use for some reason :)
The great thing that just using `logging` module is enough to log something like `Downloading file A, 55% [55Mb/100Mb]` (absolutely equivalent in terms of information to a "graphic bar") and also happens to be composable in a way that I can then reuse that package as part of anything that is also non-interactive.
I have been using it for 4-5 years now.
A fun quirk is that tqdm.notebook didn't work with Colab's dark mode so the text was readable; this was very recently fixed.
E.g. passing an unbounded range will jus show some info about the past iterations.
The example works because range objects implement __len__ which is where tqdm will get it's total/count information from.
while True:
pass
while False:
pass
What's impossible is solving the halting problem 100% of the time with 100% accuracy. Most people don't need to do that. Solving the halting problem most of the time and then saying "I don't know, that's too hard" the rest of the time is of immense practical value, and many practical systems (the kernel's eBPF verifier, the thing in your browser that detects stuck pages, etc.) do exactly that.In this particular case, tqdm solves the halting problem in cases where it's easy: https://tqdm.github.io/docs/tqdm/#__init__
> total: int or float, optional
> The number of expected iterations. If unspecified, len(iterable) is used if possible. If float("inf") or as a last resort, only basic progress statistics are displayed (no ETA, no progressbar).