This sounds like it does peak memory, which is critical for batch jobs, since that's the bottleneck. Memory is fundamentally different than performance in that it's a limited resource, instead of cumulative cost; making any part of the program faster almost always helps speed up the program (at least a little, or at least reduces CPU load), but optimizing non-peak memory has no impact. You have to be able to identify the peak in order to reduce memory usage.
If you want peak memory profiling for Python that also runs on macOS, check out https://pythonspeed.com/fil/ (ARM support has some issues, but once I unpack my new Mac Mini I plan to fix it.)
Ways memray is better than Fil:
- Native callstacks.
- More kinds of reports, and ability to do custom post-processing of data.
- Much lower overhead (but not always, see reply).
- Subprocess support.
Fil I suspect has better flamegraphs: https://pythonspeed.com/articles/a-better-flamegraph/
And if you're running Python batch jobs, and want both peak memory and performance profiling in production, check out Sciagraph: https://pythonspeed.com/sciagraph/
(You can probably cobble together something like Sciagraph with py-spy + memray, but you won't e.g. get timeline reports designed with batch jobs in mind.)