Strict APIs vs. Forgiving APIs
publicobject.com
publicobject.com
0: Impossible to get wrong/DWIM
1: Compiler/linker won't let you get it wrong
...through
9: Read the correct LKML thread and you'll get it right - SET_MODULE_OWNER
...all the way up to 17. I'll let you discover the rest yourself :-)
But if the semantics is "substring(int start, int length)", then I prefer the partial forgiveness for the length parameter: if start + length > len(s), then assume length was actually len(s) - start; but if length < 0, still throw an exception.
Meanwhile you want to save time/work for the caller. You dont want the caller to need many lines of boilerplate just to setup the call.
public String substring(int beginIndex, int endIndex) {
int length = length();
checkBoundsBeginEnd(beginIndex, endIndex, length);
...
Where checkBoundsBeginEnd[2] does: static void checkBoundsBeginEnd(int begin, int end, int length) {
if (begin < 0 || begin > end || end > length) {
throw new StringIndexOutOfBoundsException(
"begin " + begin + ", end " + end + ", length " + length);
}
}
(as it should, in my opinion).[1] https://github.com/AdoptOpenJDK/openjdk-jdk11/blob/f0ef2826d...
[2] https://github.com/AdoptOpenJDK/openjdk-jdk11/blob/f0ef2826d...
Say you have an API that is documented to accept params A,B or C i.e /v1/api?A=1&B=2&C=3
What should happen if you pass a param D to it? i.e. /v1/api?A=1&B=2&C=3&D=4
The two most common schools of thought are:
1) Ignore D
2) Throw an error
Both present their own problems, esp. when D may be closely related to A,B,C. Interesting how API design also tends to side with personal preferences for strictness or leniency
For example, Twitter pages not working if there is a fbclid parameter.
There are two main reasons that users might validly use unknown options. The first is that it's an option with valid semantics in different versions/implementations of the protocol. The other is that you're reusing an option block for many API calls, and some of the options may not be relevant for some of those APIs. For the first use case, erroring out (or at least warning) is usually the superior solution: you don't know what it means, so you can't guarantee that you'll implement it correctly. For the latter use case, ignoring can be a safe solution.
FWIW, there's been a pushback against Postel's Law (which is what the adage you cite is usually called) in more recent times. In particular, it should be emphasized that the law is generally most applicable when you're dealing with multiple interpretations of an ambiguous specification, and is least applicable when the standards are prescribing or proscribing particular behavior.
If the first web browser was extremely strict on the tags and structure it supported then web could not have evolved that way that it did.
If you are going to make things strict at least make it hard to mess up.
they're also... unix ... timestamps.
Or what do you use to tail logs?
It doesn't store a timezone, but a UTC offset. Which is a problem because the UTC offset of a timezone may change.
This happens rarely enough that people do indeed use "ISO 8601 + offset" to store "timezoned" timestamps, but often enough that it'll probably end up corrupting your data sooner or later.
Of course, if you're storing the timestamp of a future, planned event, then yes, you need a timezone. But for historical records, store the UTC offset that was in effect when the event happened.
I don't think you can get any more clear than an ISO timestamp (WITH time zone). Unix timestamps are great and efficient, but I personally value the absolute clarity of ISO above all else.
I think it makes some sense for data formats, if you want max compatibility of independently developed software that processes the data. Not sure about APIs though. Having them fail fast on unexpected input is pretty valuable.
Granted those are two of the most successful technologies ever, so perhaps it was the right call =).
https://web.archive.org/web/20060613193727/http://diveintoma...
If you can choose whether your platform supports strict or forgiving behaviour, security interests will side with strict every time.
And indeed, HTML is no longer loose but rather strict in this protocol sense that we’re talking about, as it defines how all inputs should be parsed, leaving no scope for being liberal in what you accept. (It may surprise people, but among specs of at least moderate complexity, HTML is by far the strictest out there that I know of. I wish more specs were as strict. Actually, JavaScript is probably fairly close.)
Also from the links of dissenting opinions at the start of the dive into mark link, https://web.archive.org/web/20060616150034/http://bitworking... has a good discussion of just where Postel’s law may seem most reasonable to be applicable and inapplicable, with its two-axis (text–binary, data–language) diagram. It’s worth reading and contemplating.
That sounds like a debugging nightmare, where your production system is subtlety different in behaviour to your dev system.
Also how do you know that the forgiving behaviour is correct in production? Maybe it prevents a couple of scary error messages, but it could just as equally allow incorrect inputs to be processed and stored, creating a data cleanup nightmare later.
It depends entirely on what the rules of the API are. Function signatures in most languages lack any form of compiler enforcement of rules, so you must implement them in code, and then list the rules in the function's description. The strictness you apply doesn't matter as much as your description of what argument range is allowed, and how the behaviour is affected.
For example, substring could allow overshoot with the description "If the substring would go beyond the end of the string, the remainder of the string is returned".
What you should be concentrating on is the 80% use case of your API. What will 80% of your users need? If the lack of length overshoot support would be cumbersome to the 80%, you support overshoot. If it's useless to the 80%, you leave it out. You can also implement things as layered APIs, with more general lower level functions, and then higher level functions that are more strict. Then the 20% can use the lower level functions for their esoteric use cases, and the 80% can stick to your easy-to-use and hard-to-screw-up high level API.