228 karma · joined September 24, 2008
I'm a co-founder of the PyDataLondon monthly data science meetups and conference series, author of O'Reilly's High Performance Python.
This Friday we pull GPT apart and rebuild bits by hand. A few weeks back we tackled ARC AGI 2026. Prevoiusly we did fine tuning and making GPT funny.
It is a sort of a guided hackathon (generally I plan goals for the day) and collaborative study group.
Much fun, no money, lots of smart folk asking good questions: https://playgroup.org.uk/
Mostly that content has a scientific focus but the obvious thing that carries over to any part of Python is _profiling_ to figure out what's slow. Top tools I'd recommend are:
* https://pypi.org/project/scalene/ combined cpu+memory+gpu profiling
* https://github.com/gaogaotiantian/viztracer get a timeline of execution vs call-stack (great to discover what's happening deep inside pandas)
* my https://pypi.org/project/ipython-memory-usage/ if you're in Jupyter Notebooks (built on https://github.com/pythonprofilers/memory_profiler which sadly is unmaintained)
1. Pandas if you stay in RAM, if the team and org already know this, but learn about reduced-ram types (eg float32 rather than float64, categorical for strings and dt if low cardinality, new Arrow strings in place of default Object str). Pandas 1.5 has an experimental copy-on-write option for more predictable (but probably still not "predictable") memory usage, try to use a subset of team-agreed functions (eg merge over join) due to varied defaults that'll confuse colleagues (eg inner Vs left and other differences). Buying more ram is normally a cheap (if inelegant) fix.
2. Dask as it is an easy transition from Pandas (and it scales numpy math, arbitrary python non-math functions and lots more), lots of cloud scaling options too. Stays within Python ecosystem for reduced cognitive load. It is probably less resource efficient than Vaex/Polars
3. Ignore Dask and stick with Spark if your team already uses it, as it'll scale to larger workloads and you've taken the cognitive and engineering hit (pragmatism over purity)
Vaex and Polars are definitely interesting (hi Ritchie!), and great if you're doing research and are comfortable with potentially changing APIs but you have no legacy systems to worry about. You might buy yourself a lot of future manoeuvring room. You'll find fewer clues to tricky problems in SO than for Pandas, and have a harder time hiring experienced help.
I'm 15 years in with python and scientific work. For a lot of years I liked conda but then it got crazy slow. Next I started making conda environments and installing packages with pip. Now I'm experimenting with mamba ("fast conda") and that's pretty good.
Conda envs mean I can experiment with different versions of Python (I'm a co author for O'Reilly's High Performance Python so eg 3.11 and 3.12 are pretty interesting right now). Conda "should" also make identical teaching environments (I teach my own courses). Pip was a pragmatic choice to get installations in minutes not hours in the years when conda was silly-slow.
The above is also all for short -lived research work (my typical client mode for scientific work), so it is probably different to anyone doing long-run dev work, production deploys, or for those not needing non-Python binary support (eg GPU/C/Fortan lib support).
The scrappage ("upgrade") scheme applies to non compliant cars and vans, with a focus on low incomes, charity and small businesses, it starts in January and feels reasonably generous: https://tfl.gov.uk/modes/driving/scrappage-schemes
The scheme pays more if you replace an ICE with an EV or take a public transport pass which further supports the goal of improving London's air. As a London resident, parent and driver I'm strongly supportive of improving our air.
The geometric mean of the 3.8 to 3.11b benchmarks was a 45% speedup.
For promotion I've found that sharing appropriate ideas in a private community (e.g. data science slacks) where I'm known & trusted can be a good lead source. Tweets on particular topics help. LinkedIn helps with specific advice. Linking to relevant articles in another newsletter and getting a reciprocal mention helps (but the link has to be useful to the readers and the other audience has to care about my newsletter, else no point), that only works if you know the other author.
In each of my strategic client sessions I mention the newsletter. One chunk of value I offer is to do job listings - having your job in front of my curated audience who trust me offers much more value than e.g. LinkedIn (and a _much_ smaller reach).
In short I think try to find relevant communities, do not spam (obv.), share relevant links, offer to share something cool from someone else with no requirement for reciprocation (if it is good - share it anyway) and maybe they'll reciprocate in some way. It takes a lot of effort but is a sure fire way to build up an engaged audience who trust you.
0: https://github.com/ianozsvald/ipython_memory_usage
1: https://github.com/ianozsvald/ipython_memory_usage/blob/mast...