Show HN: Verify JSON using minimal schema
github.com
github.com
What led you to choose “!” as the “optional” modifier? My intuition would have guessed that character to have the opposite meaning.
[1]: project: https://github.com/tylerchr/jstn/ — playground: https://tylerchr.github.io/jstn/
Simple stuff like:
Person = { "name": str, "age": int }
The basic idea was to use a file full of obiwan types to document the rest api. The schema was both human-readable and valid python. The api could then trivially check the schema. Once the schema is checked, of course, its safe to go walk the json knowing that things aren't going to crash on you etc.It worked great. Still works great. I still prefer the syntax to that of MyPy.
No real adoption, of course :) But fun to remember what could have been :)
One of the benefit of JSON over XML is that it is concise and fast to work with. The standard should reflect that part as well.
With this lib, as a developer to verify a JSON for a simple REST request is as simple as - `verify(json, "{a,b,c}")`
We're counting that against products when we're analyzing what services to adopt/purchase. You doing nothing at all ... or something non-standard results in everyone else having to attempt to do it for you with the resulting inconsistencies. I'm not sure why you think OpenAPI is tailored for XML ... Here's the specification for the sample application (petstore) with service and model definitions in YAML - https://github.com/OAI/OpenAPI-Specification/blob/master/exa....
This is more for client-server JSON schema validations, comes handy when different teams are writing client and server side code. Especially in startups where iterations are rapid.
From the way you've phrased your comment, I'm going to assume that the different teams are within the same start-up ... having a versioned, executable is even better when iterations are rapid if there are breaking changes. I've been in enough start-ups to know that things are skipped over to get a product out the door. I also know that the marketing and sales guys will say "we're already using a RESTful/JSON API for our back-end, let's open our API to third-parties to accelerate adoption of our SaaS". Ouch! (but I've been there)
I will admit that "back-filling" an OpenAPI specification for an existing service can be pretty daunting ... it's quite a bit easier if you start with the first MVP and iterate the ICD with each iteration of the client and server. Admittedly, we (software developers) still have issues versioning services. It's not so bad for incremental changes but breaking changes to a service where you have to support both older and newer clients is painful.
Projects like this, mypy and others makes the whole thing bizarre to say the least.
If your program produces or consumes JSON _and_ you think it would be a good idea to ensure said JSON is what you are/a client is expecting, then you need something like this in any implementation language.
what I completely don't understand is the current fad on everything-should-be-yaml where yaml is marginally easier to write than XML, a tiny bit more than marginally easier to read (caveat emptor - Norway has a country code) and otherwise exactly as broken. (i'm talking about the 'safe' variant.)
- I was afraid it would change the language culture
- It was not comfortable to use
- The idea and design was, as you said, bizarre
- Use use a dynamic language, why the hell would I want that?
Years after the fact, I'm really happy about it.
First, most of my Python code don't use them, so Python stays python. But if I ever need them, there are here.
Of course, you could argue I should just use a language that has been designed for that since I need type. But that's missing the point: I don't usually start using hints. I code regular Python, because it's awesome. I just sometimes add type hints after the fact as a bonus.
The ability to do that is fantastic, even if the type system is far from perfect.
I want to code with Python most of the time anyway. I personally have very few use cases for another language, doing mostly web stuff.
Type hints are still very verbose, and convoluted, although it will get better with 3.9 thanks to 2 PEP targeting type hints ergonomics. So sure, it's never going to fit perfectly, but I'm glad it's here.
Now for a JSON schema, I guess it's the same deal. You use JSON because it fits your use case, and one day you realize your situation could now benefit from more robustness, and you add it to the mix.
Again, you could argue you should have though of that at the beginning, but projects don't have frozen requirements, they evolve. And often, you want to start lean, because it will most probably stays that way, and otherwise would be so costly the project would never grow to a point you would need robustness.
Now in this particular case, this lib is not just checking type, but also arbitrary logic. And making sure data is consistent is necessary no matter how dynamic or static your language is. There is a limit to the constraints your type system can represent.
My point is, I'd rather have a not so perfect one when I need it, that None at all.
And most of the time I don't need one.
What's different this and Python type hints is that this was a public-facing API, so you have far less control over what data you'll see, and the server code was Java, so getting things strongly typed and validated early on makes your life easier.
I mean the attitude that Python is "good enough" because it has type hinting and therefore it is not worth bothering with exploring other languages, especially ones that have wholly more powerful type systems. I have found (personal experience only here) that it can be tremendously hard to get Python programmers to learn and embrace other languages and approaches to software.
Type inference can get you so so far
But then again, Django as well.
I think that would be a great feature but from what I see static typing fails here, too.
Pretty much all of them. Any simple predicate like this can be encoded with witness types.
Here's an example in Java, which is hardly the paragon of static typing (i.e. it's no Haskell/Idris/Agda/Rust/Typescript):
class AlphaNumericString {
private final String str;
// use a fallible factory with a `private` constructor if you're
// morally opposed to exceptions
public AlphaNumericString(String str) throws AlphaNumericException {
if (!str.matches("^[a-zA-Z0-9]*$")) {
throw new AlphaNumericException();
}
this.str = str;
}
private static class AlphaNumericException extends Exception {
}
}
Now code can freely use `AlphaNumericString` and be guaranteed that it has been validated.You may object and say that newtype wrapping is cumbersome but:
1. That's an argument about sugar and ergonomics, not about the semantics that the static type system enforces
2. Some languages make it easier to generate forwarding methods to the underlying type (a la https://kotlinlang.org/docs/reference/delegation.html)
3. The `AlphaNumericString` is describing a smaller set of values than `String`. In general, you should be strongly considering the methods you allow and make sure that all paths continue to enforce the semantics you intend.
(s/def ::number #(<= -180 % 180))
(s/def ::my-string (s/and string? #(re-matches #"[a-zA-Z0-9]*" %)))NB Last time I used Pascal was probably 35 years ago....
Edit: I would suspect Ada would have something like this.
It would seem much more natural to make them JSON structures since they're almost that anyway.
{
"content": Required(str),
"username": Required(str, validate=username_is_ok, transform=lambda x: x.lower(),
fail_message="Username isn't valid"),
"message": Optional(str),
"some_list": Required({
"name": Required(str),
"date": Required(str)
}, is_list=True)
}
Then I can provide this in a hook to the request method. * https://github.com/keleshev/schema
* https://github.com/alecthomas/voluptuous
* https://github.com/Pylons/colanderhttps://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
I really like Elms way of Json decoding for this.
> Shotgun parsing is a programming antipattern whereby parsing and input-validating code is mixed with and spread across processing code—throwing a cloud of checks at the input [...]
> The problem is that validation-based approaches make it extremely difficult or impossible to determine if everything was actually validated up front [...]
Err, what? So validating against a well-defined schema won't necessarily cause this. But okay, again, I can buy the benefits for statically typed languages. There's more though:
> Don’t be afraid to parse data in multiple passes. Avoiding shotgun parsing just means you shouldn’t act on the input data before it’s fully parsed [...]
> Use abstract datatypes to make validators “look like” parsers. Sometimes, making an illegal state truly unrepresentable is just plain impractical given the tools Haskell provides, such as ensuring an integer is in a particular range.
I've experienced this in Java/Jackson (which btw, proves this is not exactly new, sexy, or rare in the statically typed world).
What is the suggestion for a dynamic language? In e.g. Javascript, classes will only get you so far, and seems needlessly heavyweight if you aren't going to get other benefits of type-safety. I really don't see how this is helpful, even after putting in the time to investigate this.
I want my decoders to essentially do 2 things:
1. Transform received data into a data structure that is best suited for my app. This might involve converting lists to objects, objects to sets, parse dates etc.
2. Only succeed on valid data
This way, my application never has to deal with bad data. Also, I get to design the data structures I use, not the APIs I use. By only validating and not transforming, you are pushing more advanced validation further away from where the data was received (since you'll need to transform the data at some point anyway).
As for libraries, I want composable parsers. For TS I'd probably use: https://github.com/paperhive/fefe/
and now i see lodash is a dependency.
lat = custom validator
b = shorthand for boolean (as is s for string)
! = optional
So for me reading the whole example its very easy to understand and digest.
I like how compact it turned out: https://github.com/healthchecks/healthchecks/blob/master/hc/...
Usage examples in tests: https://github.com/healthchecks/healthchecks/blob/master/hc/...
The main advantage of the approach in the article appears to be that its easy to extend validation with code - which is a nice touch. Pretty much any complex Json validation I have written has to combine both schema based validation and code based checks - usually done completely separately.
I've not used JSON Schema so far, but I think showing the simplest possible definition does not tell much about how complex or easy to use a language is.
E.g., it's easy to write C code for a working program that does nothing or just prints "Hello World". The resulting code is very simple and can be written without much experience in C.
However that doesn't make C a simple language - and indeed writing programs in C that perform useful tasks is a lot more complex and requires more knowledge about the language.
In the same way, it's nice that Json Schema has a compact syntax for the "do nothing" schema, but I think a comparison between practically used examples would be better.
(Also the "do nothing" schema would have to compare against not doing schema validation at all - and compared to that, even the two brackets are still pretty costly)
import ok from 'json-ok';
ok(24, { type: "number" });
// No issue
ok(24, { type: "string" });
// throws ValidationError: is not of a type(s) string
I wanted to use `ok-json` as a tongue-in-cheek name, but apparently since there is already a library called `okjson` npm didn't allow me to use it.