Parsix: Parse Don't Validate
github.com
github.com
Edit: Ok I think I understand...
It seems the problem would be that if you're implementing a function that takes a user email as a string, and that function is in a lower layer of the application, like inside the data access layer. It is difficult at this point to know if the email string you will be passed as input has already been validated or not. Thus you might be tempted to re-implement validation for it at your level inside this function as well and have an assertValidEmail check.
This can lead to a littering of validation throughout the code base, as each implemented function worries that the input isn't validated and re-validates it, possibly using slightly different rules each time.
Furthermore, if you decide to not validate it again, you might be left wondering, but am I sure it'll have been validated prior? How can I be sure? Someone in the future could easily start calling my function and forget to validate the email before calling it? This could eventually lead to a security issue or just a bug, by introducing a code path that doesn't ever validate the email string.
Thus if instead you'd re-write your function so it takes an email as a ValidEmail type (or object), and not as a string, you force the caller to for sure remember to validate the email first. And you also can safely assume if you're getting an email as a ValidEmail type that it has been validated. It could also technically allow you to localize the validation logic to the ValidEmail type constructor, avoiding possible duplicate attempts at validating email with different rules.
And it seems the latter "style" the author calls "Parsing" while the former they call "Validating", in the sense that since the function validating returns a modified structure it "parsed" it, because a string became a ValidEmail, thus parsing a string into a ValidEmail, as opposed to simply validating that the string is valid as an email.
And finally, this is a little library to help make use of this pattern in Kotlin.
Basically, rather than write a validation function, write a parser that returns a result of a specific type and use that type everywhere else. Then you can make sure the raw inputs are always validated.
[1] https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
If you mean "why this library", well, I guess parser combinators are nice! Some may say that a declarative statement of the parsing restrictions is better than a procedural implementation, on general principles.
As a result you can't pass a invalid (say) accountID deep into your code, bc validity is guaranteed to be checked early when you "parse" an input string into the "AccountId" type.
So: internal interfaces defined using non-primitive types, so internal methods don't need to keep validating their input. Conversion to said types happens early and predictably, catching bad values before they (eg) hit the database.
The repo presents an alternative validation approach, which is to parse the user input into a data type (or, not quite equivalently, into a class). The parsing process serves as validation. Consumer functions are written such that they only accept the parsed data type. Therefore it is now impossible for the programmer to inadvertently omit or bypass validation of user input.
The library is a set of convenience functions for actually writing these parsing / validation functions.
The rest of your program then works with this data type instead of the string and this way you will get a type error whenever you accidentally use unvalidated data.
A nice idea that goes into a similar direction is to expand on this and create more types for different levels of trust. E.g. you could have the data types ValidatedEmail, VerifiedEmail and TrustedEmail and define precisely how one becomes the other. This way your typesystem will already tell you what is valid and what is not and you can't accidental mix them up.
In this example, the user input validation step is f(String) -> ValidatedEmail, then the process of verifying it is f(ValidatedEmail) -> VerifiedEmail. But the same principle can apply to e.g. append() operation being f(List[T], T) -> NonEmptyList[T], and you can write code accepting NonEmptyList to save yourself an emptiness check. Or, take a multi-step algorithm that gets a list of users, filters them by some criterion, sorts the list, and sends these users e-mails. Type-wise, it's a flow of Users -> EligibleUsers -> SortedEligibleUsers -> ContactedEligibleUsers.
And then, why should types be singular anyway? You should be able to tag data with properties, and then filter on or transform a subtag of them. This is the area of theory I'm not familiar with yet, but I imagine you should be able to do things like:
List[User] -> List[User, NonEmpty] -> List[User[Eligible], NonEmpty] -> List[User[Eligible], NonEmpty, Sorted[Asc]] -> List[User[Contacted], Sorted[Asc]].
Or,
Email -> Email[Validated] -> Email[Validated, Verified] -> Email[Validated, Verified, Trusted].
I'm sure there's a programming language that does that, and then there's probably lots of reasons that this doesn't work in practice. I'd love to know about them, as I haven't encountered anything like it in practice, except bits and pieces of compiler code that can sometimes propagate such information "in the background", for optimization and correctness checking.
The Idris Book (https://www.manning.com/books/type-driven-development-with-i...) is structured in a very practical, "learn some then build some" format that I found to be a joy to read.
That doesn't mean you can't do it in those other languages, only that it's tedious compared to languages that are designed for it.
https://fsharpforfunandprofit.com/posts/designing-with-types...
[0] http://blog.jenkster.com/2016/06/how-elm-slays-a-ui-antipatt...
[1] https://kentcdodds.com/blog/make-impossible-states-impossibl...
But also, it’s not always wise to take this to an extreme. I’ve seen over the years many scenarios where dev teams were over-enthusiastic about this and parsed themselves into a corner by making system components over-strict and enforcing invariants that weren’t necessary to enforce, making them much harder to change later.
The right answer is, of course, somewhere in the middle, and depends on your domain and situation.
This isn’t really a criticism of the approach, it’s super useful, just that it needs to be applied judiciously. “Parse all the things” isn’t always the best advice.
>that component really shouldn’t be validating the structure of email addresses
That's right, as the article says you no longer validate data, instead you parse data. Components don't take strings and validate them, they take an EMail address.
>If you need to refine the parsing logic now you’ve got to do coordinated deployments
If you need to refine the parsing logic, you modify the constructor of the EMail class to do whatever refinements you need.
I have no doubt you may have seen codebases that did something wrong, and that wrong thing may have been related to parsing or validation, but nothing you've said indicates that there is something wrong with replacing validation with parsing, so that code always operates on semantically meaningful data types that are valid by construction, instead of opaque binary blobs that need to constantly be properly interpreted in ad-hoc ways.
> That's right, as the article says you no longer validate data, instead you parse data. Components don't take strings and validate them, they take an EMail address.
That doesn't solve the problem, by parsing it my code is caring about its syntax and structure, even if my use case doesn't require that. That creates coupling, and in a distributed system that kind of unnecessary coupling can create major headaches when trying to change things.
An even simpler example that many people can empathize with is using strict, closed enums in a distributed system (very easy to do with e.g. Java). Safely adding a new enum value is basically impossible at that point, until you make them "open" enums that can safely parse unknown values.
Again I am not disagreeing in principle, I am only saying that this has ergonomic problems in practice and is easy to misuse.
I think I would stress this point more in the README, thanks for sharing :thumbsup:
I strongly disagree. The problem you describe, of being overly strict, is orthogonal to whether you validate or parse. There is no "middle" between parsing and being strict. Parse-validate and strict-nonstrict are separate axes.
It helps reduce the risk of using invalid inputs by representing constraints over the value as part of the type.
For example: a common problem in web development security is that query parameters aren't properly validated which can lead to denial of service attacks. As a trivial example of this, consider a web server which paginates some data using "offset" and "limit" by passing those parameters directly to a database query; an attacker could set "limit" to some incredibly high value and cause the server to crash. If you're just doing validation on your inputs it's possible that some usage could end up being overlooked.
[0] https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
Does the explicit creation of a type add this introspection? I'm not convinced that it does. Now once you fix this bug, encoding it in a type prevents it from creeping into other parts of the the code. This seems more like DRY principles in action.
If there's a Limit class whose constructor and setter all check that the range is between say 5 to 100, and all existing code that needs the limit uses the instance of Limit, it just becomes less likely a code change is made that uses the limit input as it was directly provided by the user (and thus possibly out of range).
But you'd still need to have had someone be smart enough to make sure the Limit class does prevent limits that could cause DB crashes.
In practice I'm thinking, ok, so someone must have thought... Hey we should validate this user input and put in some logic for it.
So I think what this says is, validation works by having all external input validated as they are received. But it can be easy to make a code change at the boundary where you forget to add proper validation. If all existing functions in the lower layers, like in the data access layer, are designed to take a Limit object, the person who took a limit as external input and was about to pass it to the query function will get a compile error and realize... Oh I need to first parse my integer limit into a Limit, and thus reminds them to use the thing that enforces the valid range.
If instead the code had a util function called assertValidLimit, and the query function took a limit as an integer, it be easy for that person to forget to add a call to assertValidLimit when getting the limit from the user and then pass that unvalidated to the query and possibly cause a vulnerability.
And lastly, it seems they argue, if you were to validate instead in the query function itself, thus it wouldn't matter if others forget to validate since where it matters would, but then it is hard to fail at that layer, since you might have already made other changes and that can leave your state corrupted.
So basically it seems the argument is:
"It is best to validate external input at the boundary as soon as it is received, but it can be easy to forget to do so and that's dangerous. So to help you not forget, have all implementing functions take a different type then the type of the external input, which will remind people... Oh right I need to parse this thing first and in doing so assert it's valid as well.
`type Limit = 0..100;`
See discussion here: https://github.com/Microsoft/TypeScript/issues/15480
If one were only using integer types then the same problem would persist, that's correct. The problem would be solved by defining our limit type to only represent positive integers up to a specific safe value.
Type refinement is done on the input boundaries of the system during runtime to prevent errors from propagating.
I think my confusion was in trying to frame things as parsing VS validating. While I now appreciate that use of word, now that I understand, it also caused my biggest source of confusion.
That's because I think most people think of parsing as conversion, like I turn a String to an Int. Where as in your case, you're simply wanting to tag a type as having been validated, but you don't really convert the type itself, so you simply wrap it in another type in order to tag it as having been validated simply because the language offers no other way to tag the type with meta-information for the compiler to assert statically.
So because it seemed more like you're just wrapping the input, but still all code will be using the input value as it is, extracting it out of your wrapped type, the idea that you were "Parsing" and not "Validating" well just confused me.
If you wanted to go further you could start going the "correct by construction" route - having a data type that enforces more invariants. For example, you might store the recipient name, domain name, and tld of your email in separate fields. Then your parser would more obviously be a parser. I think of the "mere tag" type as a kind of degenerate case of this, rather than something totally separate.
https://andrewlock.net/using-strongly-typed-entity-ids-to-av...
The author refers to using primitives everywhere as "primitive obsession", and proposed using types instead.
https://www.markhneedham.com/blog/2009/03/10/oo-micro-types/
I think this is great first step using functional languages but you can go much much deeper than that.
https://www.slideshare.net/ScottWlaschin/the-power-of-compos...
More details: http://langsec.org/
> The Language-theoretic approach (LANGSEC) regards the Internet insecurity epidemic as a consequence of ad hoc programming of input handling at all layers of network stacks, and in other kinds of software stacks.
> LANGSEC posits that the only path to trustworthy software that takes untrusted inputs is treating all valid or expected inputs as a formal language, and the respective input-handling routines as a recognizer for that language.
Databases tend to outlive application code, or may be fronted by different applications (internal vs external for example). Keeping the constraints with the data is the best way to ensure that your data remains consistent within itself.
It looks like this library was written by someone labouring under the mistaken belief that it's better to build and use a DSL to create the illusion of declarativity than to just write a line or two of normal code (eg the focusedParse stuff).
Also, i demur somewhat at calling this parsing. It's tracking validation using typestate.
A parser combinator library can provide a good repertoire of primitives, standard interfaces for defining and using the parsers, and utilities for working with those interfaces.
User-friendly error reporting comes to mind.
If you think JSON primitives are "enough to express most input" then you missed the point; because of course you can, but no person who has learnt to use types to their advantage would want to write error-prone code and an unnecessary amount of unit tests that deal with ambiguous behavior for no good reason.
There's absolutely nothing there for "parser combinators" to do.
JSON parsing is a standalone step, and the values retrieved from it can be passed to another EXISTING (mind you) validators to handle the rest.
I live and breath this every day as I deal with forms and APIs for a living. And I never was like "oh man this is hard, if I only had parser combinators".
> If you think JSON primitives are "enough to express most input" then you missed the point
Unclear why you're trying to put words in my mouth, but JSON is enough to describe the composition of input (in terms of lists, dictionaries and leaf scalar values). Which as simple as it is, turns out to be a very significant part of the task.
Of course you have additional constraints on top of that.
But no need for "parser combinators". To JSON, a string is a string. Any additional operations performed on that string happen AFTER it's decoded from a JSON string format to an in-memory native string literal. Those are entirely independent steps. A simple function composition would do i.e. json => string => email.
And that's more or less how I parse arbitrary data input. All of which is simple to parse. Relative to, you know, say parsing a full programming language.
By no means we want to pass the message that you have to use this library to benefit from the "Parse, don't validate" pattern. What we want is to make the pattern more mainstream and provide a nice set of functionalities out of the box :)
Also I would like to encourage everyone to get as much as you need from this library. If you don't like `focusedParse`, please don't use it! xD
Allow me to disagree with the "tracking validation using typestate" bit. This is just one way, but using Parsix you can also go from a String to a more complex type like Email(local: Local, domain: Domain), where Local and Domain can also be other complex types. The point is, you should parse as much or as little as you need in your domain. That's why we call it "general-purpose parsing" ;)
It's also misleading in that the code is still doing validation, just in a different place.
In that sense, using types is only one way to do this, but you could model that in other ways. For example:
var foo = "123"
foo = validFoo(foo)
print(foo)
> {"value" : "123",
"valid?" : true}
And now you could have: function bar(validFoo) {
if (!validFoo.get("valid?"))
throw new InvalidInputException("Foo must be validated prior to calling bar.")
...
}Now types are a convenient way to do this that also gives you static checking for it, but I believe the idea is more to model that things were validated and expects validated input or fail.
That allows you to push all validation at the boundary, and make sure that no one ever forgets to validate the input, because if they do, the inner functions will fail reminding the caller: Please remember to validate this!
Stop Using Stringly Typed Data
Less boilerplate than the other solutions I know... for example it allows you to define types easily just by writing a parser/validator ("transformation" in fefe terminology).
I asked around on the work Matrix as to who actually coined it, but it's the weekend.
This is not to take anything away from @lexi_lambda, who cited her sources and documented an interesting type-theoretic approach to applying the principle. She did a great job!
If anyone wants to do a deeper dive, look into langsec, language-theoretic security. There's a lot of prior art to explore.
Meredith Patterson got back to me and attributes it to Sergey Bratus. We're a little vague on when, but it was quite some time ago.
It's a great blog post, and it popularized the slogan, which is the important part. She was quite clear in the original post to cite langsec, everyone here is on a collegial basis.
My point was not really about 'credit', it was about langsec. If the ideas in this library, and that post, are interesting, there's a lot more to discover in langsec. That's it.
So it seems like most of the value comes from standardizing on domain types like Username, Email, and so on. Using a framework doesn’t get you there, and it adds a dependency on the framework.
About missing interesting parsers, you are right, for now only the core part is done. Based on community interest, we will work on complementary packages, like more common parsers, easy integration with a web framework like ktor, effectful parsers based on coroutine, etc...
Lots of work ahead :D
The advantage is that you can write code using generics that works with any parser, but I'm a little skeptical about the value of generic code.
In the end, `Parsed` should only be handled whereever you receive the initial input, so the rest of your business logic will be free of it :)
In the end, this is just a tool and it's fine if it isn't fitting your particular style or use case ;)
So if this sort of thing takes off there is likely to be a standardization process at some point.
Go works better for this because common interfaces can be implemented by "coincidence." (Structural types.)
1) how to serialise these types into different formats (e.g. "how do I save you in a SQL database?) - inside the type, which requires one size fits all, or as separate mapping functions, which could result in significant code proliferation?
2) how to cope with partial validation, e.g. do I need a wrapper class for (say) some form inputs that have these types in its fields, but override nullable on if it's off?
3) sort of a combo of (1) and (2), but how do I serialise groups of these typed fields in different ways? E.g. I want to save a model made up of these typed fields via SQLAlchemy and I also want to publish it to RabbitMQ. How do I manage the potentially varied formats in a type-safe way? When everything was Just Strings™ it appeared to Just Work™.
4) how to compose these things with monads - e.g. if I have an optional field, should Maybe treat the value as nullable, or should the type itself contain that option for validation purposes elsewhere?
Even as I'm writing these questions I'm coming up with a potential pattern for using this sort of thing, but I'm curious to hear thoughts.
1) Just to be clear, when I read "serialise" I understand that you want to parse some object into a format suitable for storage. I would honestly go with the common flow: use some library like kotlinx.serialization or whatever your database library support. If this isn't what you meant, please let me know.
2) I personally like well defined type. For example, if you have a case where Email must be like Email(local, domain, tld) and another where you just want a String out of your email, I would rather use two different types for the two different use cases and therefore different parsing logic (which will probably share most code, but still produce different results). On the other hand, if you want to model the fact that some piece of information could be provided or not by the external source, than yes, I would just go with nullable types, which in Kotlin is like `Something?`, but in other language are known as Maybe or Optional.
3) Just to be clear, I will strongly advice against having validation in object constructors. Once this is out of the picture, you are left with simple Value Objects: simple, immutable, bundle of data with a defined structure and semantic meaning. When you work with these simple objects, it's very easy to derive different views for whatever use case you need. I'm not familiar with SQLAlchemy and Python ecosystem in general, but I guess there will be some way of converting a specific type into one that is suitable for DB consumption.
4) Both Parse and Parsed are actually Monads ;) And we know something about Monads, is that composing different Monads together is usually not straightforward xD That's said, in your particular example of having an optional field, I would see the parser produce a Maybe<Something> (or Something? in Kotlin)
Let me know if I got you right and what you think about it :)
Has anyone seen or written something like this?
I'm gravitating towards GraphQL now because strict parsing is built into it, so there is no need for all this boilerplate.
https://www.npmjs.com/package/@mojotech/json-type-validation
However, as I'm saying this, I wonder if I've been looking at this problem wrong. Since we already generate types from GraphQL schemas, maybe I should figure out how to use the same client side parser that's already in my GraphQL client, define a GraphQL schema for the types I'm interested in, and then just generate and use those types.
One thing that doesn't necessarily give me is the ability to define custom parsers corresponding to custom types. At least, I think most of that sort of thing is usually done server side with GraphQL.
So, thank you for the link and also the inspiration for considering an alternative.
I really would like to see nominal typing support in TypeScript. Currently, it's hard to validate a piece of data (or parse for that matter) once and have other functions only operate on that validated data. There are (ugly?) workarounds though [1].
[0]: https://github.com/colinhacks/zod [1]: https://gist.github.com/dcolthorp/aa21cf87d847ae9942106435bf...
You define a decoder schema, and then the resulting TS type gets automatically derived for you. You can then run data through the decoder, it will err if there's a mismatch, or return a value of the inferred type otherwise.
I honestly think the two arguments are different though. We could say that Postel's Maxim is about "how much you should parse/validate", while the problem we are addressing is rather "how you should do it" :)
Consider this. The Internet never would have become popular and inclusive like it became if it was designed to be highly strict, structured, and ordered like Signalling System 7. I like the fact that a 14 year kid can write an HTTP client or server, if they want to, and if they do it'll most likely be mostly correct. Stuff like that makes people fall in love with the Internet from the earliest ages. If we look at the history of companies like DEC vs. PC it's pretty clear technology can't survive >1 generation when there's no career track for hobbyists.
In some cases deviancy in protocols has helped standards become better. Consider HTML. If we had things Tim Berners Lee's way, we'd all be writing ugly verbose XHTML that validates. It wasn't until HTML5 that W3C embraced the chaos of the true standard and formalized the non-conformant use-cases (such as being able to omit <html>, <head>, </tr> </td>, </p>, etc.) and it made the HTML language much more pleasant.
So we should take the concerns of the "always be conservative" crowd with a grain of salt. Because they're the people who wanted XHTML. They're the people who ban numbers like 0xC0,0x80 because some version of some Oracle or Microsoft piece software had a bug once. I don't like how people who embrace policies like that invariably write code that destroys text data in some misguided effort to keep people safe from themselves. There should be better reasons for preventing things which are possible from happening.
Says lots of people. I actually wasn't even aware of that draft RFC you linked.
> we'd all be writing ugly verbose XHTML that validates
And you think that's a bad thing???!!
> It wasn't until HTML5 that W3C embraced the chaos of the true standard and formalized the non-conformant use-cases (such as being able to omit <html>, <head>, </tr> </td>, </p>, etc.) and it made the HTML language much more pleasant.
That is a crazy rewriting of history!
> I like the fact that a 14 year kid can write an HTTP client or server, if they want to, and if they do it'll most likely be mostly correct.
What I don't even... Have you seen HTTP?
In any case easy for you is clearly not a reasonable measure!
The negative consequences come later. Protocols become cultural artifacts with history. You need to study the history of bugs and 'features' to learn the de-facto standard.
To write a good HTTP client that parses and renders HTTP/HTML/CSS/ in the wild from the scratch you need to learn how Chrome, Safari interpret and implement specs.
It's possible to write future proof strict protocols with detailed instructions on how to handle new versions, unknown extensions, etc. gracefully.
I haven't seen any need to bring a library in to handle this. Though it would be nice to marry the worlds of TS, API, and Json Schema.
(You can define runtime encoder/decoders which produce typed values.)
So you either validate it all the time or you unbox it all the time.
Also with dedicated primitives you need to validate the email when you read it from the dB to hydrate it. So you removed couple of validations here but added a couple there.
My solution is stick to basic DTO and validate in the end at the service receiving the data, and let the errors, if any, propagate up the chain in a way that preserves where the input came from. In most cases I validate once.
So this issue we’re talking about. Doesn’t happen.
If you validate at constructor level, then you have this problem. However if you actually parse it, then this will have to be external to the actual class and you are left with a simple Value Object. This means that you can apply constraints only when it makes sense. Most of the time, there is little reason to validate something coming from the DB, so I would rather skip it.
About unboxing, that depends. I would expect for the overall business logic to always use the more structured types, probably this will only be needed once you need to serialize them :)
That's said, if you are happy with your current way, please keep doing it! I would still suggest to give this style a try though :)
What you ideally want is a first step that "validates", i.e. creates a representation from text that is easy to use but also succeeds for any input, and then a further "parsing" step that converts it to another representation such that the parsing only succeeds for valid input and is invertible for any value of the representation type, and then finally code that checks constraints that aren't captured by the type system.
This way you can support IDEs that need to edit partially incorrect code.
In the end, parsing is just a combination of what you described: ensure something fit a particular shape and constraints :)