Modern C++ Features – std::optional
arne-mertz.de
arne-mertz.de
My recommendation would be to just use https://github.com/martinmoene/optional-lite for cross-platform stuff.
Practically speaking, it takes a while for features implemented upstream to make it into your OS of choice. If these are really critical features, you can get third-party implementations of <optional> or you can compile your own toolchain, neither of these are particularly difficult options. It can be a bit frustrating, yes, when language feature arrives on a different schedule in the library and compiler (which happens).
Mind you, #include <optional> doesn't work on my Linux box either, with GCC. And if you look at the GCC web page, https://gcc.gnu.org/projects/cxx-status.html#cxx17, it says that GCC's support for C++17 is "experimental". I know that GCC != Clang, but let's give them some time to test things before a million projects get the code compiled in and we're stuck with the implementation.
Therefore, it's by far no longer something that's experimental IMHO. (standardised in C++17 after all)
std::optional<bool> opt(false);
assert(opt); // true even though *opt is false!
http://en.cppreference.com/w/cpp/utility/optional enum Bool
{
True,
False,
FileNotFound
};
http://thedailywtf.com/articles/What_Is_Truth_0x3f_If a developer wants a yes/no/unknown type, an enum class is better suited for the job. This forces them to explicitly compare values instead of relying on an implicit boolean conversion.
The function computing that is probably served really well by returning an optional bool.
std::optional is an useful wrapper because a simple "bool + object" would call constructors and all when that might be undesirable.
IMO, ternary logic belongs as an enum class, to require developers to be explicit about what branches are taken for what values.
I think it may be worth mentioning that std::unique_ptr can be used to "manage" non-heap pointers as well.
An optional pointer is a weird tri state thing where you can have empty or a null pointer or a non-null pointer.
There are a bunch of other parts of the C++ code base that are the same way. The most notorious one I can think of is std::declval, which just makes no sense whatsoever unless it appears inside another template.
Good luck with that.
In that case you need to pay attention to what you're doing. Semantically you'd be altering your code to convert a function that always returns a value into a function that maybe won't return it. How do you expect to replace a type with a wrapper that maybe won't wrap that type and still assume you won't need to check if a return value exists?
That would be like replacing a reference with a pointer and then forgetting to check if the pointer is null. Making this mistake has nothing to do with luck and has everything to do with incompetence, and no language specs saves you from that.
If the optional wouldn't be an implicit boolean, you'd see compile errors on the locations you forgot to change (even if you use type inference).
That isn't true, as std::optional are explicitly converted to bool to signal whether the object stores a value.
http://en.cppreference.com/w/cpp/utility/optional/operator_b...
std::optional<unsigned> opt = firstEvenNumberIn(text);
if (opt)Oh, please. Raise an exception or panic. But "undefined behavior" on dereferencing this new form of "null"? That's no good.
You say, "Oh, please." But this feature is consistent with the design mandate of the C++ language... the very foundation of its design is that you don't pay for safety if you don't want to. So I am not sure what you are trying to say.
Honestly... why aren't you just using a different language? Practically every other language invented in the past 30 years has the safety features it sounds like you want. So, why come into a discussion about some new C++ feature and complain about the fact that a safety feature is optional? This is the way it always has been for C++. This is the way pointer dereferencing works for all the new smart pointers. Dereferencing std::shared_ptr and std::unique_ptr... both unsafe, equally unsafe as std::optional. This is the way std::vector works. They're all unsafe unless you specifically use certain safer accessors like .at(). The fact that dereferencing a std::optional is unsafe is consistent with decades of changes to C++.
If I may offer a tip, when someone takes a position that seems unreasonable or baiting or argument-prone, I try to apply the principle of charity: I assume that their intentions are good.
Of course I don't always succeed at this! ;-)
If someone takes that kind of attitude, so what? You'll never lose by taking the high road; I've found it to be a very rewarding exercise to try to be chill regardless. Of course as I mentioned, I don't always succeed.
At least I try to be like the old Shadio Rack:
"You've got answers? We've got questions."
Even better would be a "safe" block like D, where everything is checked & well defined.
There are plenty of languages that make implicit, opaque tradeoffs in order to give you some measure of protection from programmer errors. C++ was never meant to be among them.
C++ is and always has been about giving the programmer tools and leaving the decision up to the programmer. If you want a language that does the choosing for you, just use something else.
I think the difference is that * and -> are normal things to do on a value. Other smart pointers also overload those operators, so it isn't an obvious red flag. For example, look at this code:
foo()->bar()
Is this safe? Depends on whether foo() returns a non-nullable smart pointer (or a raw pointer) or an std::optional. It doesn't raise any red flags when you're reading the code, and you can mess it up if you forget the signature. Sure, you could look at the signature of foo, but at that point you're looking at foo anyway and a doc comment telling you it was nullable would work just as well. (OK, it's slightly worse because you have to look at the comment. But you still have to look at the function to know what you're doing is safe.)On the other hand, in most other languages with optionals, you have something like this (Rust):
foo().unwrap().bar()
unwrap() is only used with optionals, so anyone reading the code knows this is something that might be a bug. If you forget that foo() returns an optional, you'd write foo().bar() instead, and you'd get a compiler error.The point of optionals, to me, is to have the compiler remind you to check nulls. If you can deref an optional in the same way that you deref a normal pointer, you don't get any of the benefits of optionals. (I guess it gives you a way to indicate null in places where you previously couldn't, which is useful but not why most people want optionals.)
You can get a reference to an item in a std::vector with the [] operator as you would with an array in many other languages. It does not do bounds checking, as specified in documentation. Trying to access an out of range element is undefined behavior.
Or you can use the .at() method, which will throw an exception if you go out of bounds. It is the programmer's choice.
Bounds checking is not free, same with checking for a null pointer. C++ is a performance oriented language so it gives you the choice to play it a little unsafe in the name of having full control.
I understand the idea to want the language to make it more obvious that you are doing something potentially unsafe, but it has never been the case in C or C++ that dereferencing a pointer is considered 'safe'.
You have to know what you are doing, else you might end up blowing your whole leg off [1].
[1] https://www.goodreads.com/quotes/226222-c-makes-it-easy-to-s...
Consider:
struct stuff {
int required_value;
int optional_value;
}
What would you do if you had no value for optional_value? Probably use 0 or -1 but that can be problematic, and not even an option if optional_value is not a primitive type.You could use int* optional_value instead and just assign NULL, but then you have to malloc/new and make sure to free/delete later.
You could use std::unique_ptr to handle memory management but you are still invoking dynamic memory.
std::optional allows you to have nullable variables without using pointers, even when hidden behind the scenes.
You could ask, what is the point of optional if you don't want to check if it is null? Well...I guess it's the same as no bounds checking, sometimes the programmer just knows. Most of the time, hopefully, a check will be done at some point, and subsequent accesses can skip the check.
Like it, don't like it, that's cool. But it is consistent.
https://theboostcpplibraries.com/boost.optional
https://stackoverflow.com/questions/22227839/how-to-use-boos...
"Effective Modern C++: 42 Specific Ways to Improve Your Use of C++11 and C++14" https://www.amazon.com/Effective-Modern-Specific-Ways-Improv...
"Discovering Modern C++: An Intensive Course for Scientists, Engineers, and Programmers (C++ In-Depth Series)" https://www.amazon.com/Discovering-Modern-Scientists-Program...
(Edit: Limited c++ experience, thanks for explaining)
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/p032...
Returns a value (as in std::optional) than can be null. If so you will get an error explaining what happened.
https://docs.oracle.com/javase/8/docs/api/java/util/Optional...
The Kotlin language (JVM-based) got this right: https://kotlinlang.org/docs/reference/null-safety.html
In the meantime, Java coding best practices dictate treating all references as non-null in first-party code (defensively checking references received from third-party code, such as with Objects.requireNonNull), and avoiding nullable references in favor of Optional.
The issue of Optional<T> being nullable is an extremely uncommon one - I have seen it happen once ever, and there were way bigger issues going on that happened to manifest that way. It's one of those "that's dumb but ultimately makes no difference" problems that many languages have.
You can't do for instance
struct foo {
// whatever
};
foo my_function() {
return nullptr;
}
and thus, std::optional<int> my_function() {
return nullptr;
}
does not compile (rightly). struct Vector
{
int x, y;
}
public class Program
{
public static void Main()
{
Vector v = null;
}
}
Also does not compile:"Cannot convert null to 'Vector' because it is a non-nullable value type"
Basically, just given `public foo myFunction() { ... }` there's no easy way to know
Structs are value types and cannot be heap allocated like in C++.
Doing new foo() will only call the constructor while stack allocating, while in C++ will put it on the heap.
Also, the only way structs ends on the heap is by implicit boxing when converting to an interface.
Finally, yes structs can be nullable, only when they are declared as such, Typename?.
Does mouse hovering work when you're reviewing a diff on gitlab?
Unless you can magically understand it is a struct just by looking at the name, instead of a typedef, using definition or even a macro.
So the value of an optional is either a valid value or it has no value. The default constructed std::optional<T> has no value
Also, in c++ if you try to use an empty value in an optional<T> it isn't using undefined behavior as dereferencing a null pointer would, plus with optional you get value semantics so things like copy are deep vs copying the pointer. Not sure about C# nullables, but I guess if you copy the .value member it is deep.
In C++ you can have pointers to objects, but they're explicitly denoted by an asterisk. The C++ equivalent would be "Object* myobject;"
However, you can also directly have object variables in C++. That's not something that exists in Java. A variable declared like, "Object myobject;" can never be null. The myobject variable is a value, just like i in "int i = 5;" is a value.
In C++ a useful heuristic could be to think of values like Java primitives.
A C++ integer takes up some amount of space. Maybe it's four bytes on your platform. Integers on your platform take up four bytes. Not "four bytes plus a pointer", but four bytes and that's it.
In those four bytes you can store 2^32 values. There's nowhere else to store whether the integer is null -- if you wanted to signal to someone that the value was "missing", you would either have to,
- Come up with your own "in-band" signal, like "-(2^31) means 'missing'", or
- Have an out-of-band signal, like a separate boolean variable, or keep a pointer to "integer or null". (Pointers being nullable is arguably kinda weird still.)
Functions that return an integer in C++ on your platform will put four bytes "in the right place" on the stack. (Or in a register? I don't really know what calling conventions look like...)
Again, there's nowhere to signal that the value is missing -- if you don't fill in that memory location or that register, well, there will just be some other (undefined) value sitting in its place that the caller will read. Ditto functions that expect integer arguments etc.
Now, this has some interesting consequences: for stack-allocated values in C++, "object identity" tends to be preserved less often than in Java. When you put an object into a `vector`, you'll probably put a copy there. When you mutate that copy or call some function on it, you're not mutating the original. You can get back to "Java semantics" of object identity and nullability etc by using pointers and explicit heap storage.
On a register. I don't know of any mainstream platforms which return simple values like an integer through the stack.
For larger values (larger than a word or a couple of words), on the other hand, usually the calling function passes the called function a pointer to a location on the stack where it should write the result.
Either way, like you said "there's nowhere to signal that the value is missing".
Object& obj = other_obj;
For this reason, when passing objects by reference, references should be preferred to pointers where possible. Consequently, I would translate Java code like: void f(Object obj) { ... }
Object obj = ...
f(obj);
Into C++ code like: void f(Object& obj) { ... }
Object& obj = ...
f(obj);
(Setting aside various other concerns like move semantics and const-correctness.)If for example you have a function with a return type of string, it must return a string. It cannot return nullptr, NULL, or 0. It could return an empty string. But in some cases you might want to distinguish between empty string and no result, in that case you would want to the return type to be std::optional<std::string> .
In database terms, the need for null values means that the data schema is denormalized. With a simpler approach there is seldom any need for these "not applicable" values.
And that's the problem with all languages with special support for "nullable" values: They glorify bad data structure design. I've seen codebases where 30-50% of the lines is just handling of null values that SHOULD NOT BE THERE and where nothing meaningful (or even functional) happens when the value is actually null.
Goes well together with OOP: In the name of isolation, each object is abstracted from all context necessary to construct a straightforward and correct program. The result is more and more meaningless boilerplate and burnt out developers.
Personally I'm fine with not assigning any value to "not applicable" variables. Or putting a sentinel there, like -1 or NULL. But adding a physical case (i.e. changing the type) to support "not applicable" cases is just not a good idea.
And nobody ever use a fully normalized schema in practice, because it's cumbersome and gets in the way of solving the problem.
And that's for databases, that are optimized stable storage and data coherence, not for performance. A fully normalized schema is a performance disaster if applied to running applications.
Anything beyond 3rd normal form brings very sporadic gains to maintainability and space optimization. The usual case is that you lose those. About performance, I can imagine there exist situations where you'll gain some by going further than the 3rd normal form, but I don't believe anybody sees that as a normal situation.
It's kind of like the dream of having a sufficiently smart compiler. In theory, a relational database could operate in normal form and give you all the performance you need. In practice, they don't always do that.
Are you saying it would be better for performance and space to have all of those optional fields in separate tables, and that it wouldn't be a huge PITA to write queries against for say reporting and while debugging?
I'll assume "no" since they're so many of them and they're individually optional fields. Maybe most of them are strings and you could simply use empty strings. Or whatever, it doesn't matter. I'm not saying these cases are trivial, but they're leaf concerns and not really relevant from a "schema" point of view.
But if the answer is really "yes", then I'll say that's at least a maintainability problem (bad coherence).
Our UI would have to join in all the fields, as it displays them all in one window, and our customers would absolutely hate to split filling them up into multiple windows.
Also, once the form is sent and approved by officials, the data has to be saved without change for 10 years, like saving the paper version of the form for 10 years in a filing cabinet. It also needs to be easily accessible for the users up to 3 years after approval in case one needs to make a correction, which is essentially done using a special copy of the form.
I'm no db expert by any stretch, I've learned by doing. So not gonna claim this couldn't have been handled better.
Flat tables work surprisingly well. Parent-pointers instead of child-lists work surprisingly well and bring a lot of benefits including modularity/coherence and cache efficiency. All it takes is rethinking the control flow. It's usually possible to iterate tables separately instead of walking trees (the hierarchy mindset).
Child lists or other structures that don't have natural ids can for example get "indexed" with first- and last-indices like this:
struct Parent {
some data;
Child childFirst;
Child childLast;
};
struct Child {
some childdata;
Parent parent;
};
struct Parent *parent;
struct Child *child;
int parentCnt;
int childCnt;
These indices are trivial to setup: SORT(child, childCnt, compare_child_by_parent);
for (int i = 0; i < parentCnt; i++) {
parent[i].childFirst = childCnt;
parent[i].childLast = 0;
}
for (int i = 0; i < childCnt; i++) {
if (parent[child[i].parent].childFirst > i)
parent[child[i].parent].childFirst = i;
if (parent[child[i].parent].childLast < i + 1)
parent[child[i].parent].childLast = i + 1;
}
The important thing to see is that most operations are inside whole-table loops. Not in the O(n) sense - each step does something meaningful. Nested loops or complicated data manipulations are drastically decreased compared to OOP approaches.