HNHacker News
TopNewBestAskShowJobs

jammycrisp

69 karma · joined September 11, 2014

submissionscomments
jammycrisp··on DotDict: A simple Python library to make chained attributes possible
Yep! Both msgspec (https://jcristharif.com/msgspec/supported-types.html#datacla...) and orjson support encoding dataclasses to JSON natively.
jammycrisp··on Show HN: CodSpeed – Continuous Performance Measurement
> we measure the number of instructions and memory/cache accesses through CPU instrumentation performed with Valgrind. This approach gives repeatable and consistent results that couldn’t be obtained with a time based statistical approach, especially in extremely noisy CI and cloud environments.

This is a neat approach! I'm curious how well it maps to actual perf degradations though. Valgrind models an old CPU with a more naive branch predictor. For low-level branch-y code (say a native JSON parser), I'd be curious how well valgrind's simulated numbers map to real world measurements?

My probably naive intuition guesses that some low-level branchy code that valgrind thinks may be slower may run fine on a modern CPU (better branch predictor, deeper cache hierarchy). I'd expect false negatives to be rarer though - if valgrind thinks it's faster it probably is? What's your experience been like here?

jammycrisp··on FastAPI 0.100.0 release notes
starlite was the original name, it was recently renamed to litestar due to comments about how easily confused "starlette" and "starlite" are.
jammycrisp··on FastAPI 0.100.0 release notes
+1 for litestar[1]. The higher bus-factor is nice, and I like that they're working to embrace a wider set of technologies than just pydantic. The framework currently lets you model objects using msgspec[2] (they actually use msgspec for all serialization), pydantic, or attrs[3], and the upcoming release adds some new mechanisms for handling additional types. I really appreciate the flexibility in modeling APIs; not everything fits well into a pydantic shaped box.

[1]: https://litestar.dev/

[2]: https://github.com/jcrist/msgspec

[3]: https://www.attrs.org/en/stable/

jammycrisp··on FastAPI 0.100.0 release notes
If you like cattrs, you _might_ be interested in trying out my msgspec library [1].

It works out-of-the-box with attrs objects (as well as its own faster `Struct` types), while being ~10-15x faster than cattrs for encoding/decoding/validating JSON. The hope is it's easy to integrate msgspec with other tools (like attrs!) rather than forcing the user to rewrite code to fit the new validation/serialization framework. It may not fit every use case, but if msgspec works for you it should be generally an order-of-magnitude faster than other Python options.

[1]: https://github.com/jcrist/msgspec

</blatant-evangelism>

jammycrisp··on FastAPI 0.100.0 release notes
> Maybe it was very slow before

That is at least partly the case. I maintain msgspec[1], another Python JSON validation library. Pydantic V1 was ~100x slower at encoding/decoding/validating JSON than msgspec, which was more a testament to Pydantic's performance issues than msgspec's speed. Pydantic V2 is definitely faster than V1, but it's still ~10x slower than msgspec, and up to 2x slower than other pure-python implementations like mashumaro.

Recent benchmark here: https://gist.github.com/jcrist/d62f450594164d284fbea957fd48b...

[1]: https://github.com/jcrist/msgspec

jammycrisp··on Pydantic 2.0
While it's definitely much faster than pydantic V1 (which is a huge accomplishment!), it's still not exactly what I'd call "fast".

I maintain msgspec (https://github.com/jcrist/msgspec), a serialization/validation library which provides similar functionality to pydantic. Recent benchmarks of pydantic V2 against msgspec show msgspec is still 15-30x faster at JSON encoding, and 6-15x faster at JSON decoding/validating.

Benchmark (and conversation with Samuel) here: https://gist.github.com/jcrist/d62f450594164d284fbea957fd48b...

This is not to diminish the work of the pydantic team! For many users pydantic will be more than fast enough, and is definitely a more feature-filled tool. It's a good library, and people will be happy using it! But pydantic is not the only tool in this space, and rubbing some rust on it doesn't necessarily make it "fast".

jammycrisp··on Pydantic V2 leverages Rust's Superpowers [video]
Thanks, glad you like it!
jammycrisp··on Pydantic V2 leverages Rust's Superpowers [video]
While I agree that there are ways to write a faster validation library in python, there are also benefits to moving the logic to native code.

msgspec[1] is another parsing/validation library, written in C. It's on average 50-80x faster than pydantic for parsing and validating JSON [2]. This speedup is only possible because we make use of native code, letting us parse JSON directly and efficiently into the proper python types, removing any unnecessary allocations.

It's my understanding that pydantic V2 currently doesn't do this (they still have some unnecessary intermediate allocations during parsing), but having the validation logic already in compiled code makes integrating this with the parser theoretically possible later on. With the logic in python this efficiency gain wouldn't be possible.

[1]: https://github.com/jcrist/msgspec

[2]: https://jcristharif.com/msgspec/benchmarks.html#benchmark-sc...

jammycrisp··on Pydantic V2 rewritten in Rust is 5-50x faster than Pydantic V1
It looks like pydantic-core is distributing musllinux wheels, which should work fine on alpine. Fwiw tooling like cibuildwheel makes building and publishing wheels for all the common platforms fairly straightforward now.
jammycrisp··on Pydantic V2 rewritten in Rust is 5-50x faster than Pydantic V1
Are there any necessary features that you've found missing in msgspec?

One of the design goals for msgspec (besides much higher performance) was simpler usage. Fewer concepts to wrap your head around, fewer config options to learn about. I personally find pydantic's kitchen sink approach means sometimes I have a hard time understanding what a model will do with a given json structure. IMO the serialization/validation part of your code shouldn't be the most complicated part.

jammycrisp··on Show HN: Faster FastAPI with simdjson and io_uring on Linux 5.19
If you're primarily targeting Python as an application layer, you may also want to check out my msgspec library[1]. All the perf benefits of e.g. yyjson, but with schema validation like pydantic. It regularly benchmarks[2] as the fastest JSON library for Python. Much of the overhead of decoding JSON -> Python comes from the python layer, and msgspec employs every trick I know to minimize that overhead. </sales pitch>

[1]: https://github.com/jcrist/msgspec

[2]: https://github.com/TkTech/json_benchmark

jammycrisp··on I_suck_and_my_tests_are_order_dependent
Pytest has an equally deprecating option for a different "use case":

    disable_test_id_escaping_and_forfeit_all_rights_to_community_support = True

https://docs.pytest.org/en/6.2.x/parametrize.html#pytest-mar...
jammycrisp··on Data Classification: Does Python still have a need for class without dataclass?
Integrating `dataclass` more into the language builtins might be nice, if only that it may allow/encourage a more native and performant implementation. Using a dataclass right now results in slightly slower class operations than handwritten types, and much slower import times.

I maintain another dataclass-like library[1] that's written fully as a C extension. Moving this code to C means these types are typically 5-10x faster for common operations[2]. It'd be nice if the builtin dataclasses were equally performant.

[1]: https://github.com/jcrist/msgspec

[2]: https://jcristharif.com/msgspec/benchmarks.html#benchmark-st...

jammycrisp··on DuckDB – An in-process SQL OLAP database management system
You might be interested in checking out Ibis (https://ibis-project.org/). It provides a dataframe-like API, abstracting over many common execution engines (duckdb, postgres, bigquery, spark, ...). Ibis wrapping duckdb has pretty much replaced pandas as my tool of choice for local data analysis. All the performance of duckdb with all the ergonomics of a dataframe API. (disclaimer: I contribute to Ibis for work).
jammycrisp··on Crafting container images without Dockerfiles
For creating images without docker from conda/mamba environments, there's also the existing `conda-docker` tool https://github.com/conda-incubator/conda-docker.
jammycrisp··on Porth, It's Like Forth but in Python
There's also https://github.com/llllllllll/phorth, an implementation of forth that compiles to cpython bytecode
jammycrisp··on Show HN: Robyn – A fast, extensible async Python web server with a Rust runtime
For my own "fast" projects[1] I've taken to providing benchmarks, but adding a big 'ol caveat at the top describing ways in which the benchmark may not reflect reality. Some people really want to see benchmarks! Others are turned off by projects that overreach on performance claims. I'm hopeful that adding context around benchmarks to tamper excitement strikes the right balance.

[1]: https://jcristharif.com/msgspec/benchmarks.html

jammycrisp··on The domain for the Python Requests library is expired
GitHub issue: https://github.com/psf/requests/issues/6140
jammycrisp··on The fastest tool for querying large JSON files is written in Python (benchmark)
> I should mention that spyql leverages orjson, which has a considerable impact on performance

Even with orjson, you're still paying the cost of creating a new PyObject for every node in the JSON blob. orjson is well engineered (as is the backing serde-json decoder), but any JSON decoder that isn't using naive algorithms is mostly bound by the cost of creating PyObjects. Allocating in Python is _slow_.

I wrote a quick benchmark (https://gist.github.com/jcrist/de29815389eaed4eaf5b24fbcfdab...) showing a handwritten query that accesses only a few fields in a 13 MiB JSON file. The same query is repeated with a number of different Python JSON libraries. Results:

    $ python bench_repodata_query.py 
    msgspec: 45.018014032393694 ms
    simdjson: 61.94157397840172 ms
    orjson: 105.34720402210951 ms
    ujson: 121.9699690118432 ms
    json: 113.79130696877837 ms
While `orjson`, is faster than `ujson`/`json` here, it's only ~6% faster (in this benchmark). `simdjson` and `msgspec` (my library, see https://jcristharif.com/msgspec/) are much faster due to them avoiding creating PyObjects for fields that are never used.

If spyql's query engine can determine the fields it will access statically before processing, you might find using `msgspec` for JSON gives a nice speedup (it'll also type check the JSON if you know the type of each field). If this information isn't known though, you may find using `pysimdjson` (https://pysimdjson.tkte.ch/) gives an easy speed boost, as it should be more of a drop-in for `orjson`.

jammycrisp··on Build an Elixir Redis Server that’s faster than HTTP
No use in changing things if it's working for you, but you might be interested in trying out msgspec (https://jcristharif.com/msgspec/) in tino instead of using msgpack-python. The msgpack encoder/decoder is faster, and it comes with structured type validation (like pydantic) for no extra overhead (https://jcristharif.com/msgspec/benchmarks.html).
jammycrisp··on Namedtuple in a Post-Dataclasses World
Yes, but it's also less flexible. Tradeoffs.
jammycrisp··on Namedtuple in a Post-Dataclasses World
Y'all may be interested in a fast dataclass-like library I maintain called msgspec (https://jcristharif.com/msgspec/) that provides many of the benefits of dataclasses (mutable, type declarations), but with speedy performance. The objects are mainly meant to be used for (de)serialization (currently only msgpack is supported, but JSON support is in the works), with native type validation (think a faster pydantic).

Mirroring the author's initialization benchmark:

    In [1]: import msgspec

    In [2]: from typing import NamedTuple

    In [3]: class Point(msgspec.Struct):
    ...:     x: int
    ...:     y: int
    ...: 

    In [4]: class PointNT(NamedTuple):
    ...:     x: int
    ...:     y: int
    ...: 

    In [5]: %timeit Point(1, 2)
    48.4 ns ± 0.195 ns per loop (mean ± std. dev. of 7 runs, 10000000 loops each)

    In [6]: %timeit PointNT(1, 2)
    185 ns ± 0.851 ns per loop (mean ± std. dev. of 7 runs, 10000000 loops each)
jammycrisp··on Imp: A full-stack relational language built around incremental maintenance
The author used to work on Eve.
jammycrisp··on Implementation plan for speeding up CPython
Agreed on optimizing core objects. I recently wrote a C base class (https://jcristharif.com/quickle/#structs-and-enums) for defining dataclass-like-types that's noticeably faster (~5-10x) to init/copy/serialize/compare than other options (dataclasses, pydantic, namedtuples...). For some applications I write this has a non-negligible performance impact, without requiring deep interpreter changes. Using the base class is nice - my application objects are still defined in normal python code, but all the heavy lifting is done in the c-extension.

However, this speedup comes at the cost of being less dynamic. I'm not sure how much more optimized core python objects could be without sacrificing some of the dynamism some programs rely on. Python dicts are already pretty optimized as is.