>If you are doing the validation inside of a constructor, you are still doing validation instead of parsing.
Why that would be considered validation rather than parsing?
From the original post:
>Consider: what is a parser? Really, a parser is just a function that consumes less-structured input and produces more-structured output.
That's the key idea to me.
A parser enforces checks on an input and produces an output. And if you define an output type that's distinct from the input type, you allow the type system "preserve" the fact that the data passed a parser at some point in its life.
But again, I don't know Haskell, so I'm interested to know if I'm misunderstanding Lexi Lambda's post.
The idea is that your parsed representation and serializer are likely produce a much smaller and more predictable set of values than may pass the validator.
As an example there was a network control plane outage in GCP because the Java frontend validated an IP address then stored it (as a string) in the database. The C++ network control plane then crashed because the IP address actually contained non-ASCII "digits" that Java with its Unicode support accepted.
If instead the address was parsed into 4 or 8 integers and was reserialized before being written to the DB this outage wouldn't have happened. The parsing was still probably more lax than it should have been, but at least the value written to the DB was valid.
In this case it was funny Unicode, but it could be as simple as 1.2.3.04 vs 1.2.3.4. By parsing then re-serializing you are going to produce the more canonical and expected form.
But yes usually you do want to split something into it's elemental components, should it have any.
But even with that understanding and from re-reading the post, that seems to be an extra safety measure rather than the essence of the idea.
Going back to my original example of parsing a Username and verifying that it doesn't contain any illegal characters, how does a parser convert a string into a more direct representation of a username without using a string internally? Or if you're parsing an uint8 into a type that logically must be between 1 and 100, what's the internal type that you parse it into that isn't a uint8?
Just for the sake of example, your internal representation might start from 0, and you just add 1 whenever you output it.
Your internal type might also not be a uint8. Eg in Python you would probably just use their default type for integers, which supports arbitrarily big numbers. (Not because you need arbitrarily big numbers, but just because that's the default.)
IP address would be about the minimum amount of structure. Something else would be like processing API requests. You can take the incoming JSON and fully parse it as much as possible, rather than just validate it is as expected (for example drop unknown fields)