HNHacker News
TopNewBestAskShowJobs

itamarst

5,611 karma · joined February 24, 2016

I write about speeding up Python software development, and code, at https://pythonspeed.com

I write about fundamental engineering skills and programmer career advice at https://codewithoutrules.com

I also write a weekly email about all the mistakes I've made both coding and in my career over the past 20 years, so that you can learn and avoid them: https://softwareclown.com

submissionscomments
itamarst··on Faster binary search: from compiled code to mechanical sympathy
When this was posted to lobsters someone shared this relevant link: https://curiouscoding.nl/posts/static-search-tree/
itamarst··on Faster binary search: from compiled code to mechanical sympathy
The bucket boundaries are often chosen from a random sample of the data, if the input data is very large. Sorting is O(nlogn), but using binary search per value to assign a bucket is O(n), plus the cost of creating the buckets on a sample. So once you hit a large enough number of values this scales better.

Binary search does give random access but in this case there's only up to 255 buckets typically, so it's random access on cached memory.

itamarst··on Faster binary search: from compiled code to mechanical sympathy
Better guesses reduce the number of guesses, so there will be less branch misprediction, but there will still be mispredictions for each remaining branch. So I would guess branchless interpolation search would still help.

In practice because the real code in scikit-learn is used in parallel, memory bandwidth starts being a problem in real usage. Plus, in the overall algorithm (this is just a small part) the time spent on binary search is now low enough that there are other, more significant bottlenecks elsewhere. So in practice the branchless optimization had enough impact on the original motivating code base that there didn't seem much point spending more time on it.

itamarst··on Loading Pydantic models from JSON without running out of memory
That's great! Would also be cool (separately from Pydantic use case) to add jiter backend to ijson.
itamarst··on Loading Pydantic models from JSON without running out of memory
https://remarkjs.com/
itamarst··on Loading Pydantic models from JSON without running out of memory
I talk about this more explicitly in the PyCon talk (https://pythonspeed.com/pycon2025/slides/ - video soon) though that's not specifically about Pydantic, but basically:

1. Inefficient parser implementation. It's just... very easy to allocate way too much memory if you don't think about large-scale documents, and very difficult to measure. Common problem with many (but not all) JSON parsers.

2. CPython in-memory representation is large compared to compiled languages. So e.g. 4-digit integer is 5-6 bytes in JSON, 8 in Rust if you do i64, 25ish in CPython. An empty dictionary is 64 bytes.

itamarst··on Loading Pydantic models from JSON without running out of memory
Once you switch to ijson it will not save any memory, no, because ijson essentially uses zero memory for the parsing. You're just left with the in-memory representation.
itamarst··on Loading Pydantic models from JSON without running out of memory
msgspec is much more memory efficient out of the box, yes. Also quite fast.
itamarst··on Loading Pydantic models from JSON without running out of memory
You can't add extra attributes that weren't part of the original dataclass definition:

  >>> from dataclasses import dataclass
  >>> @dataclass
  ... class C: pass
  ... 
  >>> C().x = 1
  >>> @dataclass(slots=True)
  ... class D: pass
  ... 
  >>> D().x = 1
  Traceback (most recent call last):
    File "<python-input-4>", line 1, in <module>
      D().x = 1
      ^^^^^
  AttributeError: 'D' object has no attribute 'x' and no __dict__ for setting new attributes
Most of the time this is not a thing you actually need to do.
itamarst··on Loading Pydantic models from JSON without running out of memory
For my PyCon 2025 talk I did this. Video isn't up yet, but slides are here: https://pythonspeed.com/pycon2025/slides/

The linked-from-original-article ijson article was the inspiration for the talk: https://pythonspeed.com/articles/json-memory-streaming/

itamarst··on Poireau: A Sampling Allocation Debugger
One key question in these sort of things is how free() works: it is given a pointer, and it has to decide whether this was sampled or not, with _minimum_ effort.

Poireau does this, IIRC, by putting the pointers it sampled in a different memory address.

Sciagraph (https://sciagraph.com), a profiler I created, uses allocation size. If an allocation is chosen for sampling the profiler makes sure its size is at least 16KiB. Then free() will assume that any allocation 16KiB or larger is sampled. This may not be true, it might be false positive, but it means you don't have to do anything beyond malloc_usable_size() if you have free() on lots and lots of small allocations. A previous iteration used alignment as a heuristic, so that's another option.

itamarst··on When should you upgrade to Python 3.13?
It's somewhat domain specific. Pure Python libraries have easier time supporting new releases than libraries that rely on C APIs, and even slower are those that deal with less stable implementation details like bytecode (e.g. Numba). But definitely getting better and faster every release.
itamarst··on It's time to stop using Python 3.8
It's much faster! There's been significant performance improvements since 3.8.
itamarst··on Not just Nvidia: GPU programming that runs everywhere
There's overhead in transferring data from CPU to GPU and back. I'm not sure how this works with internal GPUs, though, insofar as RAM is shared.

In general, though, as I understand it (not a GPU programmer) you want to pass data to the GPU, have it do a lot of operations, and only then pass it back. Doing one tiny operation isn't worth it.

itamarst··on Profiling your Numba code
You can use Numba to speed up some Pandas calculations: https://pandas.pydata.org/docs/user_guide/enhancingperf.html...
itamarst··on How many CPU cores can you use in parallel?
NumPy default is that you iterate over the earlier dimensions first.

The slow code is likely at least partially slow due to branch misprediction (this is specific to my CPU, not true on CPUs with AVX-512), see https://pythonspeed.com/articles/speeding-up-numba/ where I use `perf stat` to get branch misprediction numbers on similar code.

With SIMD disabled there's also a clear difference in IPC, I believe.

The bigger picture though is that the goal of this article is not to demonstrate speeding up code, it's to ask about level of parallelism given unchanging code. Obviously all things being equal you'll do better if you can make your code faster, but code does get deployed, and when it's deployed you need to choose parallelism levels, regardless of how good the code is.

itamarst··on How many CPU cores can you use in parallel?
Neat! Unfortunately at the moment it still ignores cgroups, it's just a wrapper around sched_getaffinity().

https://github.com/python/cpython/blob/6a69b80d1b1f3987fcec3...

itamarst··on How many CPU cores can you use in parallel?
That's not it. I updated the article with an experiment of processing 5 items at a time. The fast function doing 5 images at a time is slower than the slow function doing 1 image at a time (24*5 > 90).

If your theory was correct, we would expect the optimal number of threads for the fast function processing 5 images at a time to be similar to that of the slow function processing 1 image at a time.

In fact, the optimal threads in this case (5 images at a time) was 20 for slow function, 10 for fast function, so essentially the same as the original setup.

itamarst··on DoorDash raises minimum pay to $29.93 per hour in NYC
Visiting NYC a few weeks ago the bike-based restaurant delivery people were interesting to see, not a thing where I live.

Relevant to "active time", they seemed to spend a lot of time during the day just waiting around...

itamarst··on Speeding up Cython with SIMD
There's presumably a reason they've spent the past 20 years adding additional instructions to CPUs, yeah :) And a large part of the Python ecosystem just ignores all of them. (NumPy has a bunch of SIMD with function-level dispatch, and they add more over time.)
itamarst··on Speeding up Cython with SIMD
Author here. The original article I was going to write was about using newer instruction sets, but then I discovered it doesn't even use original SSE instructions by default, so I wrote this instead.

Eventually I'll write that other article; I've been wondering if it's possible to have infrastructure to support both modern and old CPUs in Python libraries without doing runtime dispatch on the C level, so this may involve some coding if I have time.

itamarst··on Some reasons to avoid Cython
Author here: Note that this hasn't yet been updated for Cython 3, which does fix or improve some of these (but not the fundamental limitation that you're stuck with C or C++).
itamarst··on How to use a Python multiprocessing module
Reminder that on Linux multiprocessing is broken out of the box, and you need to configure it to not freeze your process at random: https://pythonspeed.com/articles/python-multiprocessing/

(This will be fixed in future Python versions, 3.14 maybe.)

itamarst··on Cython 3.0 Released
Cython is really great at a small scale, but it has issues scaling up to large codebases (of Cython), where I think it's an anti-pattern. None of these are the fault of the Cython creators, it's really an extremely useful tool, it's just inherent in the design space, the things that make it so cool (transparently mix C and Python!) come with trade-offs.

1. Memory unsafety. It's still C or C++ in the end, the more you shift your code in that direction the easier it is to screw up.

2. Two compiler passes: first Cython->C, then C->machine code. This means some errors only get caught in second pass, when it's much harder to match back to the original code. Extra bad when using C++. Perhaps Cython 3 made this better, but it's a very hard problem to solve.

3. Lack of tooling. IDE support, linting, autoformatting... it's all much less extensive than alternatives.

4. Python only. Polars is written in Rust, so you can use it in Rust and Python, and there's work on JavaScript and R bindings. Large Cython code bases are Python only, which makes them less useful.

Long version: https://pythonspeed.com/articles/cython-limitations/

itamarst··on Electric bike, stupid love of my life
There are electric folding bikes you can take inside with you.
itamarst··on FunctionTrace: Graphical Python Profiler
It's nice to see how many different approaches to profiling there are these days in Python. I work on another (commercial but with free plan) Python profiler, Sciagraph: https://sciagraph.com.

The main use case is data science and other long-running batch jobs. Some differences:

1. It does memory profiling at basically no performance overhead; sounds like for FunctionTrace it's high overhead so off by default. And it catches _all_ memory allocations, not just Python API ones. This is based on using sampling, so it's not useful for profiling tiny functions (but for data science/scientific computing it'll work just fine).

2. Uses sampling for performance profiling, unlike FunctionTrace. Again, perfectly fine for any non-micro-benchmark data science program.

3. Also has a timeline view, without having to upload your data anywhere.

4. No native stacks yet.

5. Shows you if you're using CPU or I/O for every particular sample.

itamarst··on Understanding CPUs can help speed up Numba and NumPy code
Yeah, it was low again.
itamarst··on Hippo Meat Nearly Became an American Food Staple
There's a fun pair of alternate history novellas based on this: https://publishing.tor.com/americanhippo-sarahgailey/9781250...
itamarst··on 'Tylenol Lite' – Will a safer new useless painkiller replace a dangerous old one
Just so you know who is funding this site: https://en.wikipedia.org/wiki/American_Council_on_Science_an...
itamarst··on 'The People's Hospital' treats uninsured and undocumented
In normal sufficently-rich countries they call this a "hospital".
Page 1 of 34Next →