Std::visit is everything wrong with modern C++
bitbashing.io
bitbashing.io
> it’s completely bonkers to expect the average user to build an overloaded callable object with recursive templates just to see if the thing they’re looking at holds an int or a string.
You don't have to: http://en.cppreference.com/w/cpp/utility/variant/holds_alter... (and http://en.cppreference.com/w/cpp/utility/variant/get to access the value).
(get_if at least nets you type-safe unwrapping similar to `if let` in Swift or Rust, though it returns a pointer rather than a reference)
FWIW, std::get is type-safe in that you cannot specify a type outside of the variant types. It's safe at runtime in that it will throw std::bad_variant_access if the active object doesn't match the type.
And because of this, you think we may as well use enum+union? Even if you only plan on manually type switching on a variant, std::variant saves you from a lot of boilerplate.
The above parts of the standard library don't help in achieving this goal.
For example, consider a parser that matches on tokens. If you add a new token, the match should fail, because you want to guarantee, at compile-time, that every possible case is handled.
This is one reason that the lack of sum types in Go is so painful, to the point that someone wrote a special library for it [1].
If the visitor approach is acceptable but has_alternative is not, get<> also accepts index values (as numeric template parameters, e.g., get<0>(v), get<1>(v), etc.) and variant has a method called index() to give a numeric value saying what the current type is. This is easy enough to use in a switch statement:
std::variant<int, double, std::string> v;
...
switch (v.index()) {
case 0:
std::printf("%d\n", std::get<0>(v));
break;
case 1:
std::printf("%f\n", std::get<1>(v));
break;
case 2:
std::puts(std::get<2>(v).c_str());
break;
}
In this case you can add a default: branch and throw an error for unhandled types, or you can hope that the compiler issues a warning about a missing case (index is constexpr, so it's possible for the compiler to know that you've missed something). It's not as good as a compile error, but it might be good enough.Historically, C++ used includes so that it could be compatible with C. In the future, modules can be used which avoid many of the problems with includes [0].
[0] http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/n468...
> Historically, C++ used includes so that it could be compatible with C.
I think it is more the case that C++ slowly splintered off C and never broke free completely of #include. Rust is also quite compatible with C without supporting anything like #include. It even has a module system!
Sum types don't just store values of different types. They store different states, with associated data. So, for instance, consider the following simplistic expression AST; how would you store it in a std::variant?
enum Expr {
Number(usize),
Negate(Expr),
Add(Expr, Expr),
Sub(Expr, Expr),
Let(String, Expr, Expr),
Var(String),
} using NumberExpr = int;
using VarExpr = std::string;
struct AddExpr;
using Expr = std::variant<NumberExpr, AddExpr, VarExpr>;
struct AddExpr {
std::unique_ptr<Expr> a;
std::unique_ptr<Expr> b;
}
Of course, this being C++, you need forward declarations and a firm grasp of the rules of incomplete types to be confident about declaring a simple AST type.Requiring this kind of wrapping is awkward compared to e.g. Rust or Haskell's treatment of sum types, which unlike C++17 and std::visit both have powerful pattern matching features built into the language. Saying this as someone who writes C++ all day: std::visit and std::variant are weaksauce.
http://www.boost.org/doc/libs/1_65_1/doc/html/variant/tutori...
boost::variant<
int,
boost::recursive_wrapper<binary_op<add>>,
boost::recursive_wrapper<binary_op<sub>>
> expression;
When the visitor is applied, the recursive wrappers are unwrapped transparently.That's exactly what a tagged union does. std::variant is a tagged union. I was not aware of the name 'sum type' but it's supposed to be a synonim for a tagged union. Guess not, but std::variant is not meant as what you describe[1].
[1]: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/n454...
Not exactly. std::variant is one particular type of a tagged union, where the only discriminant is the type. You can also have a tagged union where the discriminant determines some semantic state, and multiple such states may store the same type of value. That's still a sum type, still a tagged union, and not something std::variant can do.
In it's terminology, std::variant is a 'discriminated union' instead of a 'sum type'. Making std::variant the latter would also have had disadvantages it seems.
Aren't these the same thing?
The reason why it's the basis, is because it's a time-tested, proven and stable solution that has been around since 2002.
The reason why it's so ugly, is because that's the best you could do in C++ back in 2002.
So, there's a perfectly rational explanation for all this - it's not "insane". It is unfortunate that they didn't come up with a better API that would make use of new language features, but it's not like someone deliberately set down to design the more convoluted older API just to confuse people.
Edit: Let's use the same example.
#define SETTINGS \
X(string, str) \
X(int, num) \
X(bool, b)
struct Setting {
union {
#define X(type, name) type name;
SETTINGS
#undef X
};
enum Type {
#define X(type, ...) t_ ## type,
SETTINGS
#undef X
};
Type tag;
};
Printing settings like in the example becomes this: void printSettings(const Setting& s) {
switch(s.tag) {
case t_string: printf("A string: %s\n", s.str.c_str()); break;
case t_int: printf("An integer: %d\n", s.num); break;
case t_bool: printf("A boolean: %d\n", s.b); break;
}
}
We can also load more things into the x-macro, so it's possible to define the switch cases above just like in the structure definition. We could add a third parameter called full_name: void printSettings(const Setting& s) {
switch(s.tag) {
#define X(type, name, full_name) \
case t_ ## type: std::cout << "A " full_name ":" << s.name << "\n";
#undef X
}
}
Disclaimer: I have not compiled or run any of this code. case t_int: printf("An integer: %d\n", s.num); break;
case t_bool: printf("An boolean: %d\n", s.num); break;
it will just break at runtime, unlike std::variant which prevents this at compile timeEdit: In fact it is possible to handle the cases explicitly while enforcing types by having the X-macros call functions in the switch/case and declaring prototypes via the macros:
#define X(type, ...) void printSettings_ ## name(type);
SETTINGS
#undef X
void printSettings(const Setting& s) {
switch(s.tag) {
#define X(type, name) \
case t_ ## type: printSettings_ ## name(s.name); break;
SETTINGS
#undef X
}
}
Then you define printSettings_... and have them do stuff with the data.Whoa, let's not say things we can't take back ;)
Thank you for the condescension. But if that was your point, you could have written it more clearly.
Is this not a relatively direct import from boost:
http://www.boost.org/doc/libs/1_64_0/doc/html/variant.html
I remember that being relatively painful to use 10 years ago.The std::variant has a bit nicer API compared to boost with e.g holds_alternative. I expect I would create a make_visitor wrapper myself or use an open source one if needed, but it wasn't needed.
C++ is a massive improvement over straight C in pretty much every practical way. And best part is if you don't want to use awkward, heavyweight abstractions like std::variant then you can just not use them and be no worse off for it.
C++ definitely has some improvements over C, but it definitely does not have a "massive improvement over straight C in pretty much every practical way". It's pretty bad that you have to avoid many parts of the language.
I've also noticed from various projects that C projects tend to have less code to grok than C++ projects while achieving the same thing. Who would've thought.
This new stuff actually comes from boost (like most new stuff in C++)
https://github.com/mapbox/variant
It has a very handy "match" method that addresses the issue raised in this blog post.
They've also compiled many links to other variant implementations, as well as a bibliography on the standardisation efforts:
https://github.com/mapbox/variant/blob/master/doc/other_impl...
https://github.com/mapbox/variant/blob/master/doc/standards_...
> Is it expected to be common knowledge for your everyday programmer?
> (And if the goal of adding variant to the standard library isn’t to make it a tool for the masses, shouldn’t it be?)
These are great questions for any piece of software. Especially features being introduced in a language.
Teach new languages to new programmers. Then when they have firm grasp on these concepts, and for some cruel trick of fate they have to do C++ development, then they can look this up.
std::variant<string, int, bool> v;
...
if (auto pstr = std::get_if<string>(&v)) {
cout << *pstr << endl;
}
else if (auto pint = std::get_if<int>(&v)) {
cout << *pint << endl;
}
else if (auto pbool = std::get_if<bool>(&v)) {
cout << *pbool << endl;
}
else {
count << "null" << endl;
}
which is not much worse than the pattern matching syntax.Surely a good thing.
But I agree, the struct seemed fine (not amazing, but good enough) to me.
Among other things, it's a design with weaker coupling.
Perhaps they expect Boost to work round it until the next standard is released...
I think maintainers of C++ have a case of "functional programming" envy.
In object oriented languages this pattern is mostly unnecessary, because inheritance is usually better solution for the same problem. Only reason to use something like this in C++ is (maybe even only perceived) efficiency gained by removing level of pointer indirection.
It really comes from functional languages like ML, and is extremely useful and convenient when combined with pattern matching. In fact, attempts to emulate it in OOP using interfaces and visitor pattern tend to bring a lot of boilerplate, and obscure the actual logic, which is exactly what the author is complaining about.
There's definitely more reason to use it than just "efficiency gain", which is why many languages introduced in the last decade have it built in (e.g. Rust, Scala or Swift).
Could you say something about the difference between the two? I'm not aware what it is.
Edit: wow I accidentally a word.
data Choices = Good | Bad | Ugly
like I can do in haskell, where a value of type `Choices` can only ever be one of those options. But there doesn't seem to be any reasonable approximation in python. class AbstractChoice:
# defines all common operations on choices
# plus potentially some useful stuff on top of those
pass
class Good(AbstractChoice):
# conatins the code specific to good things
pass
class Bad(AbstractChoice):
pass
class Ugly(AbstractChoice):
pass
This isn't a straight-on replacement though. It's especially bad when you want to separate your concerns not along the good-bad-ugly-axis but something different (which is when you'd e.g. go on to use mixins, or that ugly visitor pattern we've seen in the OP).If you want to make a closed set of subtypes that's usually possible only by writing a class hierarchy with private subtypes and an abstract outer type with factory methods for the inner types. That takes hundreds of lines for even just a couple of variants, and then there is still no compiler help for exhaustive switch/match.
FP doesn't have many things C++/D etc. have. Should we insist on every FP to have assembly-level access, custom memory managers etc. because they are cool as well?
More importantly it's the second category of type relations. There are product types and there are sum types, if you only have product types, you're missing an entire half of expressible type relations/compositions.
It's not a matter of functional versus imperative, there's nothing inherently functional about sum types.
In OO it's "abstract class PaymentMethod" and then another 100 lines, plus probably a horrible visitor pattern (because you don't want the payment method handling the payment). This to just define the type, no actual payment logic.
Huge amounts of boilerplate, and very little power if the compiler can't ensure exhaustive matching.
It's already possible to make the class hierarchy in most/all OO languages but it takes 150 lines for what can be expressed in 5 with some help from the language. Most importantly, without lang support the compiler can't validate exhaustive matching.
yes. all the languages should converge (and the relevant ones actually do)
std::variant<std::string,bool> v {"abc"};
Q: Which type is used ?(Since I'm asking a question, you know there is a gotcha). Answer: "abc" is const char*, which can be converted to bool and is picked because the way variant is designed.
I also understand (and strongly support) the committee's desire not to introduce yet another language construct when a library solution can be worked out, given the horrible beast c++[11,14,17,20] has already become.
More generally, I write numerical software and I do with there was a better option for writing this software outside of C++, but I don't see one right now. Specifically, C++ gives us direct access to the c-api of other languages and some pretty powerful tools to handle that. As such, if we want our software to work across multiple languages like Python, MATLAB/Octave, or a variety of other languages, C++ appears to be the best fit. Yes, it's possible to hook something like a Python code to MATLAB/Octave, but it's hard because Python and MATLAB/Octave handle memory in different ways. For example, they differ on how and when objects are collected by the garbage collector, so it makes it hard to use a Python object directly in MATLAB/Octave. In C++, we have enough tools to handle hooking C++ objects and items to other languages. Certainly, it's a pain, but I contend it's easier than to hook two other languages together through the c-api, but I find this more difficult to manage and we now have a bunch of additional code to maintain as well. As such, as many disadvantages as the language has, I appreciate new features like std::visit because it means that I can write easier code for my algorithms and still be able to hook to others.
Finite amount of committee time. It is better to ship working pieces and add missing pieces later than wait forever for the perfect solution. std::variant had been in bike-shed-mode for a very long time due to the never-empty-guarantee saga. At soon as some sort of consensus was reached, it was decided decided to ship what was ready.
I've edited my original comment to fix the error.
This means that (ignoring the mis-dating of the unique/shared_ptr) he already admitted, his comment is otherwise accurate.
He explicitly said his reference was C++03
If these are causing hangups, you are overengineering your code.
bool batch_failed = std::any_of(
jobs.begin(), jobs.end(),
[](const auto & job) { return job.failed(); }
);
In C++03-plus-auto, you'd have to define an helper function or class: bool isJobFailed(const Job& job) {
return job.failed();
}
And then use find_if, I guess: auto last = jobs.end();
auto iter = std::find_if(
jobs.begin(), jobs.end(),
isJobFailed
);
bool batch_failed = (iter == last); jobs.Any(job => job.failed())
(Lambdas in C++ is one ridiculous example where I must use all the existing types of brackets in one expression: [](){}
is a lambda.)What humanrebar wrote is an algorithm that will work with any range of any type that has a 'failed()' member function.
Here's another formalisation:
auto batch_failed = [](auto const& batch) {
return std::any_of (begin (batch), end (batch),
[](auto& testcase) { return testcase.failed (); });
}
This function will basically work with anything.99% of the time, I work on the full container, why not providing an overload which let me write at least
bool batch_failed = std::any_of(jobs,
[](const auto & job) { return job.failed();}
);
Then, lambas are verbose.
I would like to have a simpler syntax like for simple lambas (no capture, single expression in the lamba body). job could be automatically typed with const auto&, or you could write it yourself if you wish. job => return job.failed();
And a the current one, which is more verbose, for more complex lambas (capture, several expressions in the body)and then there's things that are wonderful to use even if it's terrifying to look at how it's implemented like std::forward which is used with vector.emplace_back.
there's also simple things like vector<unique_ptr<X>> being legal C++11 syntax instead of an illegal right shift operator.
I think my current GCC 7.2 should support it, but I don't think my current Clang 4.0.1 does. EDIT: Just tested and I can confirm that.
[1] http://en.cppreference.com/w/cpp/utility/variant/visit#Examp...
switch(setting->tag) {
case Str:
// do something with setting->str
break;
case Int: // do something with setting->n
break;
case Bool: // do something with setting->b
break;
}"""
Here be dragons, though, since we must always remember to:
* Update tag whenever assigning a new value.
* Only retrieve the correct type from the union (according to tag).
* Call constructors and destructors at appropriate times for all non-trivial types. (string is the only one here, but you could imagine similar scenarios with others.)
"""
...the last paragraph is the part that will almost surely trip a C++ newcomer up if they tried this out. The first to are (valid in my experience) complaints about maintenance mistakes that result in compiling programs but runtime errors.
The big problem with C++ is that the template fanatics took over. Templates are a crappy programming language - bad syntax, confusing semantics, and tough debugging. But there's no way to stop people from extending the language via templates. Hence Boost, and "you are not supposed to understand this" templates.
(I fear that Rust is going down the same rathole.)
Sum type has been the standard functional programming language term for this concept since at least the 1970s, perhaps earlier.
std::variant<size_t, std::string> var = "test";
std::visit([](auto &&val) { std::cout << val; }, var);
If you also want to print the name of the type of the variable, along with the value itself, well, C++ doesn't have reflection (yet). So you're going to have to write a function that takes a value and returns its type as a string, which you can already do using the typeid operator. This is basically implementing reflection yourself, it is cumberstone but doesn't require advanced template programming.There no need for any of the madness with explicit types inside the visitor lambda, which is deemed as neccesary by the author.
Edit: here's another way to map types to strings. So no templates necessary at all for this entire problem.
std::map<std::type_index, std::string> typeIdxToString;
typeIdxToString[typeid(std::string)] = "string";
typeIdxToString[typeid(int)] = "int";
// ...
Edit2: Turns out C++ has some form of reflection after all. You can just use: typeid(val).name()
If your compiler supports pretty names.These questions seem utterly opaque to me with modern C++. Perhaps I just need to spend a weekend reading the docs again?
std::variant will never allocate by itself.
sizeof(std::variant<char, uint8_t>) == 2
so there is no padding between the values and the discriminanteg. look here: https://godbolt.org/g/AdLAi1 there's not a single new / malloc / whatever
That's what I'm getting at, there's this whole new suite of functionality that is sort of opaque to me, whereas I'm used to having a grasp of what's going on at-a-glance when working with C and C++98. I'll just have to spend a weekend working with it.
I'm pretty sure that it's not a matter of compiler brightness; the C and C++ standards are both pretty explicit about when the automatic and dynamic storages are used. If anything, I guess that it would be harder for compilers to allocate such things on the heap.
I checked and tcc doesn't allocate anything for instance for this code; neither does gcc 4.4 in c++98 mode :
int main()
{
struct foo {
int a, b, c;
char x[3000];
} f;
}Then the answer is yes. We haven't advanced to the point where the compiler can guess that you want. If you can give me an actual example of when it would be an insurmountable task to do so, please let me know.
That, and that it should be that hard, and the provided features for doing so should be better, is the entire point of the article.
>We haven't advanced to the point where the compiler can guess that you want.
That's not some case of magic compilers. This is just bad design.
>of when it would be an insurmountable task to do so, please let me know.
Whoooosh. The whole point is not that it is insurmountable, but that it's much worse than it should be.
Also, your visit function does not do exactly the same thing as the author's example. The author's example also prints the type's name. How would you do that without making the visitor cases explicit?
See my edit.
std::visit([](auto &&val) {
std::cout << typeIdxToString[typeid(val)] << val;
}, var);
> E.g. you would want the '+' operator to do different things for strings and numbers.Then overload the '+' operator, no need to put all that in the visitor lambda.
Edit: Replaced second "useful" with "elegant"
printf is a good, easy example of a case where you need to have different logic depending on which variant you are. It's not an endorsement of printf qua printf over cout qua cout. cout only works here because the standard library has already overloaded cout for each variant, as the author ends up doing. If you were doing anything else, you'd need either the explicit overloading or the if-constexpr thing.
The example from the Rust book (which, full disclosure, I wrote because I <3 tagged enums and pattern-matching and the previous examples weren't that great) is probably a better one: https://doc.rust-lang.org/book/first-edition/enums.html
enum Message {
Quit,
ChangeColor(i32, i32, i32),
Move { x: i32, y: i32 },
Write(String),
}
For these four variants, your behavior is very different. What you want to be able to write is something like loop {
match get_message() {
Message::Quit => {
return;
}
Message::ChangeColor(r, g, b) => {
print!("\027[38;2;{};{};{}m", r, g, b);
}
Message::Move {x: column, y: row} => {
print!("\027[{};{}H", row, column);
}
Message::Write(str) => {
print!("{}", &str);
}
}
}
which (modulo any errors from me not actually testing this) is perfectly valid Rust code, that's entirely readable even if you only know C++ and not Rust. As far as I know, you can't write anything anywhere as straightforward as this in C++. You'd need to define at least two new classes, probably four, for the four variants, plus another function that's overloaded on the four classes to handle the behaviors; you can't have the behavior be inline in your existing function, as above. (The way I wrote this, it's just using the global print function, but it'd rapidly get messy if you needed to pass a reference to a console object or whatever.) Alternatively, you would in fact need the constexpr trickery the article suggests, which would let you write it inline with some lambdas. None of this is needed in a language with language-level support for tagged unions and pattern matching (of which Rust is hardly the only one - please don't take this as advocacy of Rust in particular); you can just write normal control structures as above.Alternatively, none of this is needed in a language with dynamic typing; you'd just do
while True:
msg = get_msg()
if msg.type == "Quit":
return
elif msg.type == "ChangeColor":
print "027[38;2;{};{};{}m".format(msg.r, msg.g, msg.b)
...
but presumably you're using C++ because you want to be able to do this without that level of dynamism (which puts some unenviable lower bounds on efficiency).This is so irrelevant as to the point of the article that it is funny.