Show HN: Pydantic – Data validation using Python 3.6 type hinting
pydantic-docs.helpmanual.io
pydantic-docs.helpmanual.io
Python comes with the jsonschema module out of the box.
It's a slight pity that the python jsonschema package doesn't support some of the more powerful recent features though.
"JSON support for named tuples, datetime and other objects, preventing ambiguity via type annotations"
https://github.com/m-click/jsontyping
If you are interested, please have a look at the first unit tests to see how it works:
https://github.com/m-click/jsontyping/blob/master/tests/test...
Note that the tests currently use the "ugly" NamedTuple syntax to be compatible with Python 3.5 and 2.7.
Does/will Pydantic handle all the standard dunder fields like __eq__, __lt__, __hash__, __cmp__ and faux-immutability like namedtuple and attrs do?
In the large, I've been learning about Rust recently and wrapping my head around the design-patterns of static typing. For internal data-structures the benefit is not as clear, but for serialising and deserialising external data (like from config files or JSON APIs) I really prefer having specific, named types instead of a generic bucket of dicts.
API documentation can be more concise. You can say "this argument must be an instance of BuildArtifact" rather than "this argument must be a dict with an 'href' key whose value is the URL to a build artifact and a 'hash' key whose value is the SHA256 of that artifact" in every relevant API.
Debugging is easier when inspecting a variable starts with "<BuildArtifact ...>" rather than just dumping a dict at you.
If you need to operate on a particular kind of data, a named class gives you an obvious place to hang a method, instead of having a loose function rattling about. For operations between two data-types (like 'merge' or 'intersection'), a loose function might still be the most appropriate, but operations like searching or summarizing are naturally methods.
In short:
* I started out with a class that substituted environment variables into itself for settings, still in aiohttp-devtools: https://github.com/aio-libs/aiohttp-devtools/blob/master/aio...
* using annotations occurred to me
* it worked
* it was 50% faster than trafaret which I'd been using before
* I added some unit tests, published to pypi and started using it.
__eq__ makes sense, I'll do it when I get round to it
__hash__ would be nice but far from simple to do in a performant way.
__lt__ I'm not sure what this would mean?
__cmp__ no longer exists in python 3.
A data validation framework is not a toy project.
When a field raises an error I've noticed the value gets replaced with a "None" type (or maybe its removed from the passed data object) when using @validates_schema. This is annoying because I have to pass the original data and see if the value is actually null or not. This can suck when using JSONAPI because you have to be super careful with your data extraction (eg. data.get('data', {}).get('relationships', {})...etc).
I would like more control over how validation is executed. The ordering of validation, the ability to stop validation or continue validation at arbitrary points. Maybe in my validator I could do something like "raise ValidationError(msg, stop_validation=True).
I would like more control over pre and post dump/load order. Sort of like a z-index in css. So when using marshmallow-jsonapi, I can specify pre_loads that access the data before and after the jsonapi pre_load formatting.
I would like a better way of using "class Meta". Right now its annoying to inherit from a base then define additional "class Meta". I think I ended up subclassing SchemaOpts and settings my defaults that way on my base schema.
I would like a way to replace error messages so instead of "Missing data for required field", it would say "Please specify data" (or whatever). I want to define this at the schema level too. I don't want to have to constantly define fields with the same behavior everywhere. It would be cool if Marshmallow had a dictionary of error codes and messages. So it would look like {'1': 'Missing data for required field.'} and I could override that error with "errors['1'] = msg".
If I had to sum it up. I like to write small functions and classes that have limited uses. A lot of times I feel like Marshmallow pushes me into more monolithic work so I can control the flow.
I use marshmallow all the time and love it - I agree that error handling is its weakest area, even using strict mode.
Because it reuses python's typing system it should have the most pythonic and flexible description of types possible.
I agree about the need for complex validation chains relating to numerous fields, that's already partially possible with pydantic (although not documented). I'll add support for this stuff as well as documentation over the next few weeks.