Don't use booleans (2019)
luu.io
luu.io
Well you knew what was coming -- Rust enums are excellent[0], and so are Haskell's[1] (try and spot the difference between an enum and a record type!)... But that probably won't help you at $DAYJOB.
A bit more on topic though -- I'd like to see a strong opinion on Option<SomeEnum> versus SomeEnum containing a Missing variant. I usually lean towards Option<SomeEnum> but I wonder if there are any devout zero value proponents.
I don't like how golang does/requires zero values, but in more expressive languages I do waver sometime between Option<T> and T w/ a "missing" indicator variant inside.
[0]: https://doc.rust-lang.org/book/ch06-01-defining-an-enum.html
function operation(user: User, state: “active” | “inactive”): void
Boom. Done. Enum defined. You want people to use enums? Remove all context switching from the definition process.def foo(bar: Literal["active", "inactive"])
Seems like there are at least 4: Mypy, Pytype, Pyright, and Pyre
https://github.com/DetachHead/basedpyright
(assuming you can't/don't use MSFT's pylance)
That said, I hate that typescript has many ways of doing it. That's the biggest problem.
There's:
enum SomeEnum {
First,
Second,
}
const enum SomeConstEnum {
A = 1,
B = A * 2,
}
type SomeEnumType = "first" | "second";
const SomeObjectThatIsAnEnum = {
"first": 1
"second": 2
} as const
I always lean towards the `enum` keyword, because I am of the opinion that if you're optimizing for less generated JS (and enums are actually worth optimizing for in this way) you're probably using the wrong language to begin with (unless you're in the browser and have no choice, etc).The reason is that the compilation from `SomeEnum` to JS involves conversion to an object indexer[0].
Should you, in the future, need to (or accidentally) remove or change an enum option, this will fail with an index out of range exception at runtime that is not obvious to track down.
Put it another way: it leaves behind an artifact in the output JS that can be a source of exceptions.
[0] https://www.typescriptlang.org/play/?ssl=5&ssc=2&pln=1&pc=1#...
How is this different from a const object style enum? Or are you more referring to the sum type enum as the better choice here to avoid this.
Generally, if you remove an enum variant, the compiler should warn you if you've used it anywhere else (though of course it can't figure out a dynamic usage)... But this seems like a problem mostly for dynamic code, where you'd probably write some sort of Enum.parse() or Enum.from() static method anyway?
If I'm understanding your comment correctly, the real problem is that when you store the value, remove one, and try to read it back out into the enum it will fail - but in the code I've written so far trying to construct the the enum object almost always consisted of using a function to validate anyway, these days I use zod but before I'd write stuff like:
function isEnum(obj: unknown): obj is Enum {
return obj && typeof obj === 'string' && Object.values(Enum).includes(obj)
}Like Typescript might (for example) represent enums as integers like 1, 2, 3... at runtime.
What happens if you decide to store a “delivery status” enum in a database? You would just be saving a number into the database, which can be hard to understand.
If it’s a string, then the value stored in the database is clearer and doesn’t depend on you being lucky with the Typescript compiler that the same enum type is compiled to the same integers after new releases.
I only realised this issue now, but it’s what I understood after reading that comment.
There may be a really good use case for the enum type, but I think in most cases, string type unions are clearer, easier, and less prone to errors.
I like Typescript's union of string literals approach because its serialised value is never in question. I know exactly what "NORTH" serialises to.
but in the code I've written so far trying to construct the the enum object almost always consisted of using a function to validate anyway
It's exactly this problem and it's not obvious for someone choosing an enum type that the underlying implementation in JS is an object indexer.The gist is: define your enum values as a const array. Then you can generate a TypeScript type based on whatever values you have in there. And, since it's an actual array, you can easily write guards / asserts.
So you end up with something easy to extend and also has compile-time and run-time checks.
Otherwise, it becomes an absolute mess of “almost” defined correctly.
This comparison appears to be unintentional because the types '"a"' and '"b"' have no overlap.(2367)
https://www.typescriptlang.org/play/?#code/MYewdgzgLgBAZiEAu...It's features like these that make Python/Ruby/etc. type-checkers insufficient substitutes for TypeScript.
I think my example wasn't the best. Of course "active" | "inactive" would get reused. But if it's just a little function flag it could easily be a one-off.
The advantage of Enum with Missing variant is that it's more expressive, and that you don't have to nest access in the Some matching.
Generally I prefer Option<Enum>, and often you can match directly with Some(Variant) to avoid nested matches.
Strictly speaking, I think providing "map" just makes it functorial. Monadic would need a flatmap. (In addition to the other functor and monad requirements, of course.)
pure :: Monad m => a -> M a
(>>=) :: Monad m => M a -> (a -> M b) -> M b
pure is sometimes called return and >>= is sometimes called bind and flat map.I was unsure that bind was the same as flatmap but it looks like it is.
For example, the Option and Result type both have functions like "map", they do the same thing just on different types. They're not quite generic in that sense, but on a high level they seem so.
Another example are reactive libraries. Everything is pegged into some common operations like map, take, and so on.
Java $DAYJOB here. If I'm a caller, I probably won't realise that SomeEnum.Missing exists. I'll just see that a function accepts a SomeEnum (which I don't have) so I won't call it.
Funnily enough, Java 1.8 is what really made me start taking Haskell and languages with better type systems more seriously -- seeing that Option actually deserved to be basically everywhere broke my frame.
The nice thing about a class is that when some bizarre edge case comes up (and whose Extract, Transform, Load process has not?) one can isolate the weirdness.
It has the same problem as Booleans by supporting only two branches: single happy path and single unhappy path, when real life may have many paths, and it’s not always clear which of them are happy or unhappy (from the business logic point of view).
Even for purely technical errors, simple OK | ERROR isn’t enough, but should be something like OK | RETRY | UKNOWN_ERROR | FAILED.
fetch(accountId, history = true, details = false);
However, I would be opposed if this mechanism is permitted to perturb the order of fixed arguments. If history is the third argument, rather than the second, it should error.I.e. "history = true" means "we are passing true as the second parameter, which we believe to be called 'history', on penalty of error".
Fixed parameters having their names mentioned and checked should not be confused with the language feature of keyword parameters, which passes an unordered dictionary-like structure.
I work mostly in C++, and this is a feature I've wished for for a long time.
fetch(accountId, history: true, details: false);> "When you use named arguments in a function call, you can freely change the order that they are listed in."
That violates a key requirement in my specification, which means that I regard it as broken.
Positional parameters should take arguments only in their designated argument position.
Only non-positional parameters (keyword parameters) should be allowed to be specified in any order.
https://learn.microsoft.com/en-us/dotnet/csharp/programming-...
struct SomeArgs { bool normalize{false}; int stride{2}; char ch{'x'}; ... some other stuff };
void someFn(SomeArgs args);
someFn({ .stride = 1, .ch = 'y' }); someFn({ .normalize = true });
not as simple as python but it's close-ish and more readable than: someFn(false, 2, 'c');
That said... this article feels like a bit of a strawman. Does anyone by default use three booleans in cases that are arechetypical enums?
Then people avoid touching it when expending a bit, so they just add another boolean for the corner case they're dealing with (potentially as an optional last param). And so on, until someone refactors the whole function bearing the burden to retouch all its legacy that accumulated up to that point.
I love it when a language's bool type is just a sum-type like the sum-types you define in user code, e.g. https://ocaml.org/manual/5.2/core.html#hevea_manual10. It's an indicator of a language with good foundations.
https://existentialtype.wordpress.com/2011/03/15/boolean-bli...
typedef struct options_s {
bool toggle_case;
bool strip_whitespace;
} options_s;
char *modify_string(char *str, options_s options)
{
if (options.toggle_case) { /**/ }
if (options.strip_whitespace) { /**/ }
return str;
}
int main(void) {
char str[] = "Hacker News";
modify_string(str, (options_s){ .toggle_case = true, .strip_whitespace = false });
}
Now that I think of it it's probably trivial to forbid the use of struct literals without designated fields in code in a linterMaybe we get anonymous struct function parameter declaration with C32 ? :D
EDIT: I have been asking around to people fluent in standardese and if you leave out fields in a struct literal you are guaranteed they will be zeroed-out
modify_string(.str=str, .options = { .toggle_case = true, .strip_whitespace = false });Anyone know what the reasons have been for the language not adopting this?
It can't be anything but the most trivial thing to implement.
It's actually using unnamed parameters at all. While this issue is more frequent for boolean parameters, it can happen for integers or floating point numbers as well. We do a lot of numerics in C, and calls such as
optimize_foo (&state, 0.3f, 1e-5f, 1.f);
have the same issues as mentioned in the post: Error prone, more difficult to review, more risky to refactor.But stating “don’t use booleans” is a terrible absolutist ideas that will lead to terrible code. Imagine a student reads it, takes it to heart, and starts creating dedicated enums for every logic branch in their code.
foo(arg1: boolean, arg2: boolean, arg3: boolean)
Then when you call it it will look like foo(true, false, true) which is not great for readability. Instead I move all the arguments into an object which makes each field explicit. Ie.
foo({ arg1: true, arg2: false, arg3: true })
This also carries the advantage that you can quickly add/remove arguments without going through an arduous refactor.
The only languages it supports are JavaScript, TypeScript, PHP, and Lua. I can confirm that it works for JavaScript and TypeScript, I haven't tried the other two.
A similar extension named "Inline Parameters Extended for VSCode" supports a different set of languages: Golang, Java, Lua, PHP, and Python. It seems to be a fork of the aforementioned.
Using object arguments is a great start, and I think using enums can also be powerful (if you're using string literal types). But often times I reach for explicit interfaces in these situations. IE:
type FooScenario = { arg1: true; arg2: true; arg3: true };
type BarScenario = { arg1: false; arg2: true; arg3: false; }
type Scenario = FooScenario | BarScenario;
...etc
This provides semantic benefits in that it'll limit the input to the function (forcing callers to validate their input), and it also provides a scaffold to build useful, complete test cases.
In [1]: def my_func(first, *, second, third):
...: print(first, second, third)
...:
In [2]: my_func(1, second=2, third=3)
1 2 3
In [3]: my_func(1, 2, 3)
---------------------------------------------------------------------------
TypeError Traceback (most recent call last)
Cell In[3], line 1
----> 1 my_func(1, 2, 3)
TypeError: my_func() takes 1 positional argument but 3 were given
In [4]: def my_func_2(*, first, second, third):
...: print(first, second, third)
...:
In [5]: my_func_2(first=1, second=2, third=3)
1 2 3
In [6]: my_func_2(1, second=2, third=3)
---------------------------------------------------------------------------
TypeError Traceback (most recent call last)
Cell In[6], line 1
----> 1 my_func_2(1, second=2, third=3)
TypeError: my_func_2() takes 0 positional arguments but 1 positional argument (and 2 keyword-only arguments) were givenThe generalization of the author's lesson is the New Type idiom. Don't use basic types like "integer" when you can invent types which correspond to the actual meaning of your values. A Duration, a FileDescriptor, a DatabaseHandle, a KeyboardScanCode. Once we do this it's obvious that if the function takes a FileDescriptor and a Duration we should not give it this KeyboardScanCode we have twice, that's clearly nonsense.
Rust piles on extra value for its own New Types (as yet you cannot directly grant this to your own types although if they're enumerations they get it for free) - niches. Result<(), OwnedFd> is the same size (a single 32-bit integer) as the C int you'd have used for this purpose in a much less capable language, but since it's a Result type it has all the same ergonomics as if it was much more complicated and heavyweight than that.
I don't know why more languages don't have named parameters...
// Booleans + Named Parameters
foo(should_scrub=not options.scrubbing_disabled)
// Single-Use Enum
import ScrubbingOption, ShouldScrub;
foo(
if options.scrubbing == ScrubbingOption::Disabled {
ShouldScrub::No
} else {
ShouldScrub::Yes
}
) foo(options.scrubbing == ScrubbingOption::Disabled ? ShouldScrub::No : ShouldScrub::Yes)
it doesn't look much worse at allIf you have two-option enums and for some reason you transform them into two-option enums that represent essentially the same thing, of course that's going to be unnecessarily complex.
As a general rule, "it's more verbose" isn't much of a concern to me: verbosity is a very poor indication of complexity, and entering code with your keyboard simply isn't the bottleneck for writing code (if it is, you aren't thinking about what you're writing enough).
"For some reason" being anything so prosaic as they come from different libraries.
That doesn't fix the possibility of adding a third possibility either (what if we separate non-tax invoices into two types of invoices?).
RaiseInvoice(InvoiceType::Tax, Notifications::Email);
var taxInvoice = false
var sendEmail = true
RaiseInvoice(taxInvoice, sendEmail)
It's obvious, but people tend not to do it. RaiseInvoice(/*taxInvoice*/ false, /*sendEmail*/ true)Not being able to use logical operations on them is a feature, not a bug: logical operations on unnamed values here are going to be even less readable than passing them into functions.
function foo({bar, baz}) { ... }
foo({bar: 1, baz: false})
And correct me if I'm wrong, nothing worse than an incorrect pedant.
The ([…, ]{…}) approach is called “options object” which is a separate entity and is very convenient compared to e.g. python’s semi-infinite arguments lists.
In other words, this blog post is practically irrelevant to those languages.
My ideal language would have function prototypes that look something like:
return_type function_name(parameter_name type (default_value) [constraints], ...)
If the default value is not set then the parameter must be passed or it is a compile error. function foo({x, y, z=1}) { ... }
It eschews positional parameters though, but makes passing aptly-named variables as arguments easy: const x = 1; // Usually a complex computation instead.
const y = 2 * x;
return foo({x, y});https://www.youtube.com/watch?v=_ahvzDzKdB0
It's one of the best programming talks I've come across. For 1998, all his points on growing a new programming language are spot on.
And then Java became the opposite of a lot of the goals he enumerated.
Rules like this aren't that useful IMO. What if you have like a SetActive(bool active) method? Does this also get an enum? Should this be two methods? Blah blah blah. You can only cultivate taste with experience and consideration, not with rules. Everyone's seen codebases that followed all the rules but were still total messes. Or as Tool said: think for yourself; question authority.
I'd be inclined to do away with the argument entirely, instead going with separate SetActive() and SetInactive() methods.
> Should this be two methods? Blah blah blah.
I read about half of of submissions, I don't know if others also said all this already.
Lord knows I still do that most of the time though...
@enum MyEnum A B
sometest = x == A ? B
Which is correct but that one that I'll add new elements to the enum @enum MyEnum A B C
and forget to check if that code is still correct. Suggestions?typescript saves the day again
fetch(accountId: AccountId, options: {includeDisabled?: boolean, history?: boolean, details?: boolean})
Of course, doing this adds overhead to the developer at the time the code is written. Instead of just writing 'bool', you need to go into the header and define a new enum and decide what it and each of its entries will be named, and make sure that it is propagated up through any dependencies that will interact with that function and so on.
Frankly, that can be a lot more work - especially in a large legacy code base.
Sometimes just a simple Bool is easier.
use_boolean = false;
type boolean = (true, false, EFileNotFound)
initwithenablefan:enableAirConditioner:enableClimateControl:enableAutoMode:airCirculationMode:fanSpeedIndex:fanSpeedPercentage:relativeFanSpeedSetting:temperature:relativeTemperatureSetting:climateZone:
/s