No evidence for nudging after adjusting for publication bias
pnas.org
pnas.org
I like to provide that HN community with some context as to what this means.
There are some 300 “research” departments in each of the major social sciences: psychology, sociology, economics and anthropology. If you believe what they say, about half of their mission is teaching and the other half is research. That’s a lot, tens of billions of dollars.
The nudge findings were among the few to not only reach the level of public knowledge but, more importantly, directly influence on public policy. To use the one I most familiar with: the so called default for defined contribution retirement plans, eg 401k. These government regs assumed, for good reason, that maximizing contributions was in the public interest. Based on the nudge findings, after much debate and effort, they were updated to dictate that the max options forms was pre selected in the brief it would cause more individuals would opt for that as opposed to contributing zero.
So far so good, right? In fact nudge has become a canonical example in introductory public policy courses as to how their research can in some sense make things better.
This meta-analytic finding turns on the authors’ method for measuring publication bias. Because I accept that, I must believe that this entire body of research, probably the signal behavioral economics work, is essentially worthless! Thus, all that effort has not only been wasted but the credibility of social science in general is damaged.
Adding this to the well/known gamesmanship in peer review, debate over tenure and etc. means it’s past time to reform a large chunk of academia.
I ask because I am sure that changing defaults DEFINITELY works, especially if the user does not have a strong existing preference.
You're not really changing user behaviour most of the time, you're changing the outcome of what they're trying to do, which is to reduce their cognitive load by ignoring as much as they possibly can.
I mean I just found out two weeks ago you could change the hacker news banner color. Are you telling me I’m in a statistically insignia can’t minority of hacker news users?
Also how many settings are there in the average application, you can’t tell me most users go through all of those settings to get exactly what they want.
I guess there must be further detail in the paper and I will have to read it to understand the nuance.
Id say this is a large part of the reason Gmail, Android and Chrome exist.
Another question is whether this increases the total amount of successful donations. I was looking around for studies and found this one [1], which basically says "in some countries, yes".
[1] https://behavioralpolicy.org/wp-content/uploads/2020/01/Does...
That is, all you're doing is tricking people who didn't read carefully. People don't know they've opted in and would opt out if you called and told them that they checked the box.
I find it generally plausible that defaults don't matter much for what people consider very important decisions. I have minimal experience in this area, though.
https://blogs.bmj.com/medical-ethics/2017/09/25/organ-donati...
Capital wants to have the spigot left on. If people don't feed the beast voluntarily, Capital will make that the default.
My pet theory is, these results hinge on, “does it scale?”
Like, yes, you can do nudges and see behavioral changes. But what about when everyone is doing it constantly? Then people will get fatigued and form countermeasures.
Imagine this dynamic in another context:
“Guys, guys check this out, people are guaranteed to buy your product if you show arguments for it to random people!”
But, oops, centuries of marketing later, advertising isn’t automatically effective enough to cover its costs, people don’t automatically believe the ads.
https://sparq.stanford.edu/solutions/opt-out-policies-increa...
I grant that it would be surprising if it had no influence at all, but I think the effect is more the social signal that you should want to save the max, that your neighbors probably do (it's the default after all), etc., rather than people completely ignoring/missing it.
Again, it would be surprising if it didn't matter at all, but not unimaginable. What you're saying is that almost everybody in your company would have contributed a lesser amount if not for the default. It means you can all afford to give up $20k or whatever in income this year. There are other factors.
There is no truth to the matter of "whether defaults change behavior". This thread started about 401ks and then was taken into color preferences on the web. If someone has a gun to their head is asked if they want to die, I'm sure we'll agree that whatever the default is doesn't matter. Whether defaults do anything depends on what we're talking about. Nudges might work in web ux but not economics, why is that so incredulous?
Defaults are very strong when there is a lot of uncertainty about the payoffs of different answers.
It's a no brainer that defaults will alter outcomes for users who aren't willing or capable of making a selection for the choice in question.
Now, you can say all day long that those aren’t causal studies, but there is just no way that confounding factors like different cultures explain it, because cultures just aren’t sufficiently different, or rather cultures that are otherwise pretty similar have vastly different donation rates.
A lot of the replication crisis imo is just realizing that landmark studies were underpowered. That is, they don’t prove what they meant to prove, but that is very different from whether the effect exists i.e. an effect may exist yet be hard to prove and social scientists are rarely rigorous in study design, from training and from inherent difficulty.
Food is another example, I like cheese sometimes but when there is an option for it I take it out of the food most times but I won't go out of my way to ask for its removal otherwise, this has a real health impact.
If you’re just concerned about inflation, don’t like risk, and don’t mind locking you money up for a little bit Treasury Inflation Protected Securities [2] are also a thing. Their returns are tied to the Fed’s measurement of inflation (CPI).
1: https://www.fool.com/investing/how-to-invest/index-funds/ave...
"How can these horrible critics say nudges don't work? Have they never been nudged with a loaded revolver? Can they not imagine that working?"
Leaving those of us who don't follow the controversy closely in the field and are interested in what has actually been found out and understood about the world with some degree of confidence across many fields of study, leaves us scratching our heads unable to see through the viewing window for all the mud getting flung.
I don't think that's entirely true If anything this just highlights how complex behavioural science really is, as they're dealing with surprisingly complex humans and their surprisingly complex lives. Behavioural science is a young field.
The methodology should (I haven’t investigated theirs in detail) not be susceptible to this, and I doubt a mean of effects would make it through peer review for reasons including the ones you’ve mentioned.
This is a pretty short article, how are you confident of such a broad conclusion? What makes you that confident that this meta-analysis is decisive?
In many schools, these social science departments are a favorite for the weaker students who don't really do so well with math. They're usually filled with athletes. They love to absorb pop psych results like Amy Cuddy's Power Pose and so they don't want to listen to anyone question their results with lots of meta analysis. They want some basic ideas from class in between lots of time on the playing field.
I'm afraid that their demands will far outweigh any desire to force the fields to search for absolute truth.
I strongly disagree with this statement, even as someone who believes “nudge” effects are wildly overblown.
It means “these studies failed to find evidence” - NOT that there is nothing to find.
The distinction is important because, as it turns out, the policies that the research influenced did work, in many cases. 401k contributions did go up, in many cases. More people became organ donors. More Europeans got stronger privacy protections.
“The power of defaults” is such a cliche because, in many cases, it works.
The problem with these studies is overstating the effect - not spewing worthless BS.
That we need to “create” the idea of a “nudge effect” when it’s clear people take on commonly encountered social behaviors is bizarre.
Cognitive experience is a for loop with memory; for time spent in situation X, memory forms at rate Y. Social science solved.
Social science derives all it’s conclusions by studying the same old physical world as physical science. It’s restatement of science customized to cultural tradition. It’s cultural tradition to over hype our specialness selling books and big ideas, when the math is the same everywhere. Creating cultural objects of obvious math is a commodity now.
Perhaps due to the PR efforts of leading researchers, it was much more than “set defaults intelligently.” The interpretations were more like: we can use social science to shape peoples’ behavior at the margins. Further these marginal changes would cumulate to substantive and lasting societal improvement.
On reflection, it seems to me that the value of this paper stems from its attempt to measure or quantify publication bias. In this case, the bias was positive in the direction of with studies confirming nudge effects.
Taking that a step further implies that the actual net nudge effects across published and unpublished studies were statistically and therefore substantively insignificant. Hence the use of the term worthless, i.e. non-findings.
To say that it is costless to implement a nudge scheme in the behavioral economics sense is simply untrue. In the retirement case it required a lengthy ethical and legal debate; some study and political argument as to the best outcome, which is in part a redistributive question, hard costs associated with revision or development of messages and other materials, etc.
Worse I believe is the damage done from attention and action predicated on now seemingly faulty social science. What could’ve been done instead and what will happen in the next time a social scientist claims an ‘easy’ way to make things better are costs.
That step is in no way supported by the evidence provided.
This is not what statistical significance implies. This misunderstanding, and its inverse, leads to the very errors for which you criticize the "nudge" papers.
More to the broader point, "set defaults intelligently" in fact implies the ability to "shape peoples' behavior at the margins." Otherwise, why bother thinking about them?
That's why what is actually at issue with "nudges" is effect size & context: how much of a difference can we have, and where?
And to that question, this paper provides little insight. It aggregates too much & ignores real-world policy evidence.
Now, it's still a good paper - people have gone WAY overboard with nudges in silly places - it just needs to be understood as "let's reign in expectations" and not "this field is bunk"
Whenever I see a new term being introduced as an explanation I am hesitant to accept it, as it may turn out to just be explaining the planetary motions with epicycle, when the motions can be easier explained by moving the sun to the center of the solar system instead of the earth.
First of the definition:
> A nudge is a function of (condition I) any attempt at influencing people’s judgment, choice or behavior in a predictable way (condition a) that is motivated because of cognitive boundaries, biases, routines, and habits in individual and social decision-making posing barriers for people to perform rationally in their own self-declared interests, and which (condition b) works by making use of those boundaries, biases, routines, and habits as integral parts of such attempts.
I find this definition overly permissive and overlay reliant on unnecessary cognitive terms (like judgement and choice; which can be shortened to behavior) or economic terms (like rationality and self interest). As a fan of behaviorism this feels like an attempt to introduce epicycle into a theory that doesn’t need it. This effect—if it exists—can probably be adequately explained with good old classical conditioning and conditional reinforcements. This is the first red flag. That is not to say we can’t look for specific cognitive functions which makes some reinforcement contingencies more effective then others, but nudge feels a bit too general to actually be of any use in a model. It in fact reminds me of Albert Bandura’s theory of self-efficacy, a theory that seems to have reach a dead-end at this point.
The second red flag is the economic presuppositions. When I skim through the literature it feels like they are creating a band-aid on the thoroughly debunked notion of Homo economicus (the believe that human individuals always behave in a rational way optimized for their own self interest). So instead of recognizing the fact human behavior is more complicated, what they try to do is invent a new term to counter-act the instances where biases are “preventing” such a behavior pattern. I find such an effort to be doomed to fail, as—despite the persistence of economists—rational behavior means a different thing for each individual, and there is no “patch” for what economists call “biases”.
Umm, I have some news for you.
Further, isn't this using the same data from the original meta analysis that did find "of small to medium size" effect[0]?
Why would this, alone, undo decades of research and clear, bright-line conclusions such as the ones cited in my sibling comments? In other words, why is this letter the final word on the topic of "nudge", to you, and not the original meta-analysis? Sounds like you think everyone should pack it up and go home, all because of one letter using an alternative set of definitions and analysis.
Just seems like an overreaction on your part, especially given how vocal and... you-sounding (for lack of a better term) the "anti-nudge" crowd often is.
To take a wider view, a comment like yours is a more malicious form of nerd-sniping[1], especially on HN. Claim to have relevant credentials, voice a contrarian-but-popular-here opinion, and make a wild conclusion to give those reading it a feeling of "inside baseball."
[0] https://www.pnas.org/doi/full/10.1073/pnas.2107346118#sec-3
I don’t see how this would make me reinterpret all those successful results. Maybe I don’t understand what this is saying.
As for your own AB tests, you have seen the processes that go into them and do not need to adjust for unknown biases. So when they demonstrate a nudge effect, you can believe it.
I’ve not done any rigorous research, but I’ve participated in projects that resulted in dramatic shifts towards customers choosing what the dev team thought was the “best” outcome, just by altering wording, or making “dangerous” choices harder (such as by requiring more clicks to enable).
I assure you, Opt-outs are much less likely to happen when combined with not explaining there is a choice to be made in the first place.
Okay, you and I care about very different things.
> Okay, you and I care about very different things.
Clearly it is being suggested that it doesn't matter with respect to the social studies departments being in dire need of drastic reform. If you don't care about that either direction why are you commenting on this thread? You do actually care one way and you're "point scoring" to further the argument? Something else? I'm misreading something that I think I'm reading clearly?
>means it’s past time to reform a large chunk of academia.
And you're taking exception to it, but now just claiming you're actually not, that these one sentence "you're wrong" responses are really something much more modest.
Out of interest do you work in the field? Have ties as a graduate to one of the departments? Or are you completely disinterested when assessing the research?
If I have an interest it's that quality research is performed that advances human knowledge with some kind of efficiency of the resources spend. ie fund something that is being done properly and well to some effect over something that has been shown to be run by those who seem to be utterly incompetent or fraudulent shysters or something else that engenders zero confidence.
https://www.google.com/search?q=cass+sunstein+site%3Astatmod...
>> And you're taking exception to it
TameAntelope was addressing premises to that conclusionary comment.
If the premises do not hold, then TameAntelope does not need to address that conclusion.
But they do tho'
/tips hat to TameAntelope's style
Hitch your wagon to Cass Sustein et. al. by all means... Famously successful and successfully famous. What else do you need to know? ESP might be a thing too, nobody has proved it isn't. But we have shut down the research departments at universities involved in that BS...
https://www.google.com/search?q=cass+sunstein+site%3Astatmod...
First of all, preregistration is not a requirement for the scientific method, which has functioned well for centuries. That is a recent trend in response to the overflowing amount of haphazardly published science.
Second, it is up to the individual scientist to decide to preregister or not. Some social scientists may preregister.
Third, small sample size may be a fair critique, however that overlooks how difficult it is to collect such data.
You've made a lot of generalizations here that amount to, "social scientists aren't as rigorous as other areas of science, therefore we should only believe studies that disagree with their results". I don't think throwing the baby out with the bath water is helpful. You can take results of studies with small sample sizes with a grain of salt, watch for replication, etc. Lambasting the field as a whole doesn't make sense to me.
Finally, readers should note that this isn't a new argument. People have been making this claim about social science for 120 years, if not longer, but at least since Freud and contemporaries began publishing.
I think it’s more “ignore them completely.” It brings “science” into disrepute to let social science associate with the other sciences.
> I don't think throwing the baby out with the bath water is helpful.
There is no baby!
> People have been making this claim about social science for 120 years, if not longer, but at least since Freud and contemporaries began publishing.
Doesn’t that prove the point? It wasn’t science then and isn’t science today.
So you just want it renamed to "social studies" or what? What is your proposal, that nobody research this topic, or that they be separated in journals etc? I doubt that will have much impact on whether it makes the news. If you want that to change, you may need to get yourself onto the board of a journal you care about.
> There is no baby!
That's reductive. Just because you don't see the baby doesn't mean it doesn't exist.
> Doesn’t that prove the point? It wasn’t science then and isn’t science today.
No, it just proves it's an old disagreement, like nature vs nurture.
There is plenty of work in social science that contributes to humanity. It will always have smaller sample sizes due to the nature of collecting the data. The work can be considered useful nonetheless.
https://www.quora.com/Is-social-science-a-real-science/answe...
Humanity is observable. It's just hard to collect the data. I think you can make the case that some science isn't as rigorous, or that the jury is still out, but to say that it isn't science at all is wrong in my opinion. Even Feynman acknowledges that conclusions may be drawn later.
Aristotle philosophized quite a bit and is considered an early contributor to the scientific method. More food for thought:
https://blogs.scientificamerican.com/cross-check/is-social-s...
Or maybe people just want to shine light on the fact that social science is harder than other science for a bunch of different reasons, bring social scientists' attention towards the tools that help mitigate this, and bring the journals that seek profit over reliable results into disrepute?
Most of the papers published in current social science journals are not science, and this has been a problem for those 120 years precisely because the techniques used in chemistry or physics are inadequate for the problem domain, so applying them blindly does not produce scientific outcomes.
The comment to which I was replying lacked the nuance in yours. Context is everything.
Yes. And that the rest of us stop treating it as science, citing it as science, and relying on it as science.
For example, there is a major trend in the law of treating social sciences as having truth value the way real sciences do. That’s the kind of thing we need to stop doing.
It seems unlikely that all of social science will one day be declared as "not science". Aristotle's methods, for example, did not require a certain sample size.
You're applying far too strict of a definition to science. Basic forms of science can be practiced by a child at home. Journals publish more in-depth analyses, and it's up to them what to publish, at the risk or reward of gains and losses of readership.
[1] https://blogs.scientificamerican.com/cross-check/is-social-s...
This isn’t a fun theoretical exercise. In the public sphere, “science” tends to get invoked with dispositive weight. And it should in many cases. But for that to work, “science” must meet the level of rigor people associate with “science.” To be called “science” it should be like physics in terms of providing truth value, not psychology.
Aristotle didn’t do “science.” He was a philosopher. His ideas were precursors to science, but weren’t science.
It's possible to form a truth about human behavior, for example, I did X because Y. If someone points my behavior out to me, then in the future I may do Z in response to Y. Human behavior can change upon observation [1]. That doesn't make the study of human behavior "not science" in my opinion.
> Aristotle didn’t do “science.” He was a philosopher. His ideas were precursors to science, but weren’t science.
In that case, I suppose you will acknowledge that Galileo did science, despite not having a lot of data points. I think Aristotle did too, because I draw the line at hypothesis, observation and conclusion, which may or may not result in some ultimate truth that remains constant.
I suppose you would also say that the double slit test is science. Yet that result changes, and there is no truth value that we can explain, except by noting that the result changes when the experiment is observed.
It might be valuable and it might be worthwhile to do. But what you describe in itself is not science.
Given the definition, why wouldn't you consider observing human behavior to be science?
As has been said before, the problem is that the scientific method is right eventually. It can and often does get stuck for decades at a time, if someone with, shall we say, durable beliefs gets tenure, amasses political power and shoves their rivals out of a field. The amyloid hypothesis is just the most recent example.
Modern metascience practices (preregistration, blinding, banning "garden of forking paths" subgroup analysis, demanding high p factors and larger n) don't replace the scientific method, they're supposed to speed it up! But, by definition, these are all political issues, so they attract political arguments.
I think this is just more data. If we're all wrong for a longer time, then the impact will be more clear.
I agree that modern additions like preregistration are helpful. I only wanted to remark that it is not a prerequisite for science.
> by definition, these are all political issues, so they attract political arguments.
People are good at gaming systems. We are naturals at recognizing patterns and will adjust our behavior to meet our goals. In that sense, social science may be targeting a moving object, almost like the difference between observing and not observing the atoms in a double slit test.
As hard as it may be in social science, the process of hypothesizing, observing and forming conclusions is still science. For some, it appears that is not science because a definite conclusion never arrives.
Which viewpoint is correct? I think it's up to you to decide. And, when you don't grant people that choice, you get an anti-science response, because people naturally reject being told what to think. Science, for me, is about asking questions, not necessarily arriving at a definitive result.
It is very difficult in physics as well. Do you know how hard it is and how much effort is involved in building the LHC? Or Ligo? Or the JWST? Or ITER? They cost billions of dollars, thousands of scientists and decades to plan and make before you even get science data. Science is hard! You need to put the work and effort in, because otherwise you can't say anything about the nature of things.
But we are reforming, right? Merit based learning is over, and so really, what's it matter?
I would be shocked if that wasn't true though. Is there any evidence it's not true in that specific case, that pre-selecting the max options causes more individuals to opt for that? Have individuals opting for that indeed gone up since this was done?
Because now you have said idiots running around screaming how terrible that bias is completely neglecting the fact that everyone is subjected to it.
I'm probably totally misunderstanding, but it sounds similar to saying "there is no evidence for medicine" because you've averaged all the papers describing medical interventions that work and those that don't.
I thought the point of "nudges" is that they are so cheap to implement you can easily afford to try many. Most won't work, some will.
As for averaging, yes, you can: if a nudge is ineffective, then its result will be random, and a bunch of ineffective nudges will average zero effectiveness. The effective nudges will then push the overall average above zero. We don't see that. (The same would be true for medical interventions, unless some cause harm.)
As for being able to try lots of them: in some circumstances, maybe. But when a government is trying to nudge people towards some desired behavior (vaccination, say, to take a random example), they don't try sending out a bunch to different groups of people, then polling each of those groups to see which groups--and therefore which nudges--moved. And it's not always practical, anyway (and the vaccination example is a case where it's almost certainly not practical).
See also threads that mention A/B testing.
Ineffective and random are different. Ineffective means that the effect size is smaller than required.
For example, if you read "Paracetamol was ineffective for pain relief after surgery" it doesn't necessarily mean that the effect of paracetamol is unpredictable, inconsistent, immeasurable, or that it had no effect or negative effect. You would most likely interpret it to mean that the paracetamol did have an effect but it was insufficient - the patient was still in too much pain.
Similarly, if a nudge intervention was ineffective, it doesn't tell you how it failed, only that it didn't reach the threshold for success. And it certainly doesn't tell you anything about how well an aggregate of some effective and some ineffective results would perform.
The authors do mention that there is likely to be heterogeneity in (real) effect sizes, but somehow still go with this title/abstract.
Maybe there is a valid conclusion that some of the many nudge studies are probably claiming effects that don't exist. That could be interesting in itself. But rejecting the whole field based on this kind of argument seems wrong.
While it might be to catch the reader, I don't think it is wrong. The large problem is that we had policies implemented to nudge people in certain directions. Apart from the ethical question there needs to be hard evidence before we employ authoritarianism like that. So the headline should be pardoned, but not those that employ nudging for the time being.
We should also be wary of high-profile debunkings, now that they're increasingly in fashion due to the replication crisis and the general dour mood. It's easy to p-hack a result into significance, but you can just as easily hack results into insignificance.
These days, both findings and debunkings need a skeptical eye.
I do feel like that, even though being critical is something we should always do, that in cases where
1) the only reason you started paying attention to something was an intuitive hunch that it could matter, and
2) the only reason you started treating that hunch as established science is because you did experiments that had significant results, then
3) later you found that significance could be entirely accounted for by the file-drawer effect,
you need to adjust your expectation that there actually is an effect to lower than your expectation was at step 1). It isn't that the theory hasn't been tested (although you can argue it hasn't been tested for ingeniously enough yet), it's that it has been tested and no effect has been shown.
If you allow the existence of interest in a theory (represented in amount of ink spilled and number of experiments done) to raise your expectation that the theory is true, despite experimental indications to the contrary, you're not really doing science, you're just throwing good money after bad, probably motivated by a desire to protect the researchers and institutions that are heavily committed to the truth of the theory and/or the desire to protect other theories that depend on the one that hasn't shown results.
Are you really certain that a big debunking in PNAS, surfing a wave of other celebrated debunkings, should be taken as definitive, when a good deal of the research being debunked was published to similar fanfare in PNAS back when a different kind of research was fashionable?
I take neither the original research nor the debunking as particularly credible. Without technical expertise, I'm left to educated guess. It's just my guess.
The one that you expressed is that it "seems so obvious to you that you become suspicious." I'm just taking you at your word.
I don't even know what you're defending here other than believing your first impulse above any subsequent evidence. Nobody is preventing anyone from proving an effect, in fact they poured money into the attempt.
That's different from the existence of the phenomenon.
Same thing happened to Kahneman, Daniel (2011) and his book of Thinking, Fast and Slow. He acknowledges that several pieces of evidence he presents in the book has disappeared and can't be replicated.
He still thinks he is right, he just admits that he does not have strong evidence anymore.
What is left is a theory with less and less evidence supporting it.
The job of scientist is to find and present that evidence.
If your assumption is correct, and Kahneman failed to find good evidence that was there, that makes him a incompetent scientist. I don't think he is.
Feynman taught us to be scornful of cargo cult science, but fewer have internalized how difficult real science can be in comparison.
If this is the case, though, maybe I shouldn't be surprised that maybe no quality research has been done.
He mostly says that about just one chapter. A significant portion of the book is fallacies of basic statistics and logic.
The reality is that the way people make decisions is stupidly complex, because people have stupidly complex lives. Some tweak will work great for one project, and do nothing on the next one. It's hard to even say if it was the nudge that worked the first time.
I really view nudge theory as one of many ideas of things you can try, a tool in a toolkit. But the only tool I really feel confident works is the design-test-iterate loop.
The legal team told us we couldn't use default choices anywhere, as it could count as giving financial advice. Fair enough. So we designed the onboarding, and there was this choice the user had to make before we could create their account.
During testing, we found people were getting really stuck on this choice, to the point of giving up. The choice actually had quite low impact, but it was really technical - a lot of people just didn't understand it. Which makes sense our users weren't financial experts, which was our target user. This choice was a new concept for the market, so we couldn't relate it to other products they might know. The options inside also had quite a lot of detail when you started digging into them, detail we had to provide if somebody went looking for it. Our testers would hit this choice, get stuck, feel the urge to research this decision, get overwhelmed, give up.
We spent so long trying to reframe this choice, explaining it better in a nice succinct way, we even tried to get this feature removed entirely - but nothing stuck.
Eventually after lots of discussion with legal we were allowed to have a 'base' choice, which the user could optionally change. We tested the new design, and it made a significant difference in conversion rates.
Huzzar for nudge theory! Right? Well, maybe. I think it's a bit more complicated.
- The new design was faster. There was less screens with simpler choices. It went from a 'pick one of 5' to a 'heres the default, would you like to change it?'. Was it just the speed that made a difference?
- The user was not a financial expert, and the company behind the product was. In some sense was the user just thinking 'these guys probably know more than me I'll leave it at that'. Imagine trying to implement this exact change on something the user is an expert in - say like your meal choice in an airplane. I imagine most people would think "How rude choosing for me! I'm an expert in what I feel like eating I want to see all the options".
- It had less of a cognitive load. Like the whole onboarding flow was already really complicated, just reducing the overall mental strain to make an account may have just improved the whole experience. E.g. if we had removed decisions earlier in the flow, would this one still have been as big of an issue? We never had time to test it, so I can't say for sure.
- Lack of confusion == confidence. For the users who didn't look at the options and took the default, did they just feel more in control and confident because they weren't exposed with unfamiliar terms and choices? They never experienced the urge to research.
Like at the surface level this new design worked great, so job done. But it's hard to say definitively it was because of nudge theory. I don't think you can really blindly say "oh yeah defaults == always good" and slap them on every problem - which is why the design-test-iterate loop is so important.
> The new design was faster. There was less screens with simpler choices. It went from a 'pick one of 5' to a 'heres the default, would you like to change it?'. Was it just the speed that made a difference?
If you're just going from "pick one of 5" to "pick one of 5 but there's a default", I wouldn't expect one or the other to be "faster". Was the new design more different than that?
As for the rest, I think the beneficial features of the design are predicted by nudge theory. "Providing a credible default reduces cognitive load and confusion on the path to a decision, as the user can just trust the defaults have been set up reasonably" has always been the theory for why nudges work.
What I mean by it being faster is you could get to the next step of the process with both reading less text, and seeing less choices (just two buttons not 5). Cause if you just slapped the continue button (which most people did), you'd skip the whole explanation of all the choices.
In the context of government, a nudge means influencing people to choose something desirable, while still leaving open the option for people to choose what they want (hence preserving liberty). In contrast, a non-nudge solution would be a law or regulation that forces people into the desirable option or perhaps a tax on a certain choice.
In your UI, an example of a non-nudge solution would be removing the other options, effectively forcing their decision. Another example of a non-nudge would be charging different fees depending on their decision.
https://en.wikipedia.org/wiki/Nudge_theory#Definition_of_a_n...
"A nudge, as we will use the term, is any aspect of the choice architecture that alters people's behavior in a predictable way without forbidding any options or significantly changing their economic incentives. To count as a mere nudge, the intervention must be easy and cheap to avoid. Nudges are not mandates. Putting fruit at eye level counts as a nudge. Banning junk food does not."
I've always understood this part of the description to be more than a single one-off choice, so none of that around the decision point in the financial product would count.
E.g. speed: We could've removed earlier parts of the onboarding to make the overall experience less long, or compacted the UI so it was visually easier to skim the choices.
Expertise: we could've assured the user before the choice, that all options were good cause we're the experts and that we would've give you a bad option - so don't agonise.
Cognitive load: We could've reduced the info we showed about each option, or hidden it away behind a modal, or re-written it in plain english. The legal team told us we had to use the legal descriptions of the choices, which included technical language.
Confusion: We could've made an visualisation of the impact of their choice, that changed as they swapped between each option - showing them them a more tangible outcome of their choice. It was a complicated concept to get, so the addition of a visual aid instead of just written descriptions might've helped.
To be clear - I'd be surprised if these things would've worked, and I'm certain setting a default made a difference. The point I'm making is that I don't know for sure how much of a difference. The change to implement the default, by my eyes, also improved the overall design in these other ways as well. We didn't isolate it down to exactly what made the improvement, we were just happy it happened.
The point I'm making is you could quickly skim read this story of a team stuck on a problem, who after implementing defaults found their conversion rates jumped 11x holy shiiiiiiii- and it sounds like it's all thanks to nudge theory. It's exactly like a case study you'd see in a co-design agency's portfolio.
But in the actual real messy world of designing interfaces, it's just always a bit more complicated than that. No change is truly isolated, tested in a controlled, academic fashion. You just design your best shot each time and see what works. Because of this, it's hard to truly definitively say an improvement was because of a nudge. Best I can do is, "I mean probably" haha.
Nudges don't need to steer the user to a specific choice, just a behaviour change. Sticking with a conversion flow counts as a behaviour change.
Nudges don't need to be simple or understandable. They can be a set of complex changes where causation isn't clear. They just need to get results.
The only really hard requirement that would rule out a nudge is if you forced a choice or used financial incentives.
If you read the Nudge book you'll see that it's a political book, really. The authors introduce nudges as an alternative to hard regulation. Instead they propose that governments consider influencing behaviour in a softer way, but still leave the escape hatch open for people with strong preferences to choose what they want. This strikes a balance between state involvement and principles of liberty. (Or at least that's their argument.)
Because of this framing a nudge is defined mostly by what it isn't. It's not a nudge if it forces a user to a choice; a nudge is anything you do that changes what users do without forcing them.
This is what you've done with your series of changes that resulted in increased conversion. You've left all the choices open still, so users have as much freedom as before, but you've managed to predictably change user behaviour in a way that aligns with your goals. In other words, you've nudged them.
> However, all intervention categories and domains apart from “finance” show evidence for heterogeneity, which implies that some nudges might be effective, even when there is evidence against the mean effect.
A little bit further they say "However, all intervention categories and domains apart from “finance” show evidence for heterogeneity, which implies that some nudges might be effective, even when there is evidence against the mean effect", which makes sense. People understand stakes generally, and will likely apply different care/effort in different context, modifying the context specific effect of any given intervention.
I think the paper makes a reasonable argument:
1. There is significant publication bias in nudging studies 2. The effect of providing additional information at time of selection, or providing reminders/affirmations for self control is basically non-existent 3. The effect of modifying choice structure is indecisive. Likely we'll find that some structural modifications have strong effects in some context, but others have little or no effect is other context.
Nudges aren't just defaults. We've known for over a century that people are influenced by defaults. Nudges also aren't anchoring, where choices influence one another. Kahneman & Tversky won a Nobel prize for that and other behavioral economics ideas a decade before the idea of nudges.
Nudges are a bigger idea that many small changes lead to huge behavioral changes. Like providing a social reference point (see the average electricity use of your neighbors), surfacing hidden information (a red light to remind you to change your A/C filter), change the financial effort involved in something (deposit your drinking money into an account that you lose when you drink again; health plans that pay to stay healthy), change the physical effort of making bad choices (a motorcycle license for people who don't want to wear helmets that is much harder to get), change the consequences of options (pay a teenager $1/day to not get pregnant), provide reminders (check if an email is rude and have someone confirm they want to send it), public commitments (say you are doing X makes you more likely to do X), etc.
There are various examples of each of these working to some extent in specific circumstances.
But we have a lot of other tools for changing people's behavior. We have education campaigns. We have fines. We have taxes. We have tax breaks. The idea behind nudges is that they're an easy replacement for many of these other tools.
But the meta-analysis shows that nudges aren't a general-purpose tool that leads to significant changes in people's behavior. The behavioral changes are small, the same as we get from a fine, a tax, or an education campaign.
Aside from specific circumstances, nudges don't work better (and may be much worse) compared to our usual tools for getting people to behave.
If a category has such a broad number of phenomenon then shouldn't we be analysing individual phenomenon instead of the category as a whole? For example; defaults may work and red-light thing may not work. Why place them both under the same bucket at all? Why not study them in isolation?
And that's exactly how these meta analyses work! If you look in figure 1, they break down nudges both by the kind of intervention and by the domain. Maybe some types of nudges are much better than others. And maybe nudges work much better for say food vs finance.
Yes. Defaults have an effect, most other nudge types don't. But the domain doesn't matter much it seems.
> However, all intervention categories and domains apart from “finance” show evidence for heterogeneity, which implies that some nudges might be effective, even when there is evidence against the mean effect
So the article is saying when you look at studies of all "nudges" as whole, adjust for publication bias[0], there isn't evidence for nudges as a whole. Of course individual nudges could still have a positive impact.
Maybe I'm misinterpreting what it's saying, but as an analogy, that would be like doing a study of all "diets", determining that when you combine data for all studies on diets there isn't a positive effect, then writing an article with the claim "no evidence for diets". There's no way you could reasonably make the claim that no diets work.
[0] If you look at the studies on "nudges", studies with smaller sample sizes detected a larger positive effect size. This is because smaller studies that get a positive effect are more likely to be published than smaller studies with a negative effect. The article uses this to analyze just how strong the publication bias is and adjust for it.
Even if some studies results suggest that diets work, that does not, by itself, mean we should reject the null hypothesis (that diets don't work).
Coincidences exist: https://xkcd.com/882/
I'm fairly sure anyone who has done A/B testing at scale has plenty of evidence that nudging works. Perhaps not up to the standard of science, but there are literally people who manipulate choice architecture for a living and I'm fairly convinced a lot of that stuff actually works.
There are literally people who give astrological analyses for a living.
I think this doesn't really apply to A/B testing, because people are incentivized pay as much attention to negative results as to positive ones.
I'm sure many people here are in similar situations.
There are lots of people who do X for a living, but where X doesn't work: palm readers, fortune tellers, horoscope writers, and so on. I'm not even sure that funds managers reliably obtain results much above random.
No it's really not.
To say things a different way, I don't think this study will change anything for people actually doing choice architecture in applied settings. They have results that speak for themselves.
If folks have results that speak for themselves, then the effect more than likely is scientifically rigorously testable. It may already have been - by those very results.
"They have results that speak for themselves." Let me put my point differently. Suppose that nudges don't have any effect at all (null hypothesis). More concretely--and just to take a random number--suppose that 50% of the time when a nudge is used, the nudgees happen to behave in the direction that the nudge was intended to move them, and 50% of the time they don't move, or they move in the opposite direction. And suppose there are a number of nudgers, maybe 100. Then some nudgers will get better than random results, while others will get no result, or negative results. The former nudgers will have results that appear to speak for themselves, even if the nudges actually have no effect whatsoever.
This is the same as asking if a fair coin is tossed ten times, what is the probability that you'll get at least 7 heads. The probability of such a number of heads in a single run is ~17%. So 17% of those nudgers could be getting apparently significant results, even if their results are actually random.
This is exactly how a midwife explained to me why she uses magic crystals. She told me that there's science, and there's results, and that she's seen the crystals work.
The issue is that in fact the midwife will not have such data. The comparison being made is that A/B testing, if run competently, is pretty close to scientific research, in particular for research related to nudging.
On the other hand the dream of nudge theory is something like a study done in the UK that suggests that adding the line “most of your fellow citizens pay their taxes” will increase the likelihood that people pay taxes. This I’d be more likely to believe the benefits are not clear, and more importantly difficult to replicate across time and culture.
It seems that trying to do a meta-analysis on all of nudge theory (or large categories of it) would indeed show know impact. It’s not like you’re testing one thing, you’re comparing well designed programs, with ones that aren’t.
Lol! A/B testing in practice is rife with P-hacking and various other statistical fallacies.
If you run a useful system where it would be meaningful and interesting to know whether a social science theory actually applied, you might run an A/B test to see if it works. If it works, it is adopted—but it is almost never published. And that is for two reasons: 1. no incentive to publish and 2. major incentive not to publish. #2 is recent (post Facebook experiment) and it is specifically because a large portion of the educated public accepts invisible A/B testing but recoils with moral indignation at the use of A/B testing results in published science. Too bad: Facebook keeps testing social science theories, but no longer publishes the results.
As an example, suppose I flip a coin 1000 times and get heads 525 times. The 95% confidence interval for the probability of heads is [0.494, 0.556], so from a scientific standpoint I cannot conclude that the coin is biased. If, however, I am performing an A/B test, I would conclude that I'll bet on heads, because it is at worst equivalent to tails.
And, in opposition to your assumption: there is nothing to prevent A/B tests being published with high academic standards— like a low p value and tons of n. In an academic context, that’s just fine— it’s a small but significant effect.
A/B tests are simply controlled experiments—which are the gold standard of scientific evidence generation in psychology. My point is that the main generators of this evidence are only permitted to use this evidence to inform commerce not public knowledge. That is a loss for science and public policy, in my opinion.
Except the article is more specific and has way more details than that.
Thaler and Sunstein wrote the book on nudges, quite literally. So their definition counts, and it's the one from the article. The opt-in/out decision you mention isn't a nudge in this sense. You're not asked what you prefer, you have to be aware that you can opt-in/out and then actively pursue that option.
[1] https://www.nytimes.com/2009/09/27/business/economy/27view.h... (The issue being that in countries that are opt-out, doctors still often ask families for permission on the grounds that the deceased never made an affirmative choice to donate.)
Yes. But an alternative is to present choices on a web form without one being pre-selected but with a choice mandatory. Which is essentially what I understand Thaler to be arguing for.
>The choice has already been made.
I'd say that still is a default but one which requires more effort to change than a pre-selected option on a webform. And arguably sufficient effort that it may no longer be reasonable to default to organ donation in that manner.
Does this study imply that choice architecture plays no role in our decisions? Or am I mis-understanding it?
For example, if I search your entire house for drugs, using drug sniffing dogs and so on and I don't find any at all, that's pretty good evidence that you don't have any. It's not proof though - you might have just stashed them really well.
Similarly, if people have been looking for nudge effects for ages, doing loads of studies on it for years, and none of them have found any effect, then that's pretty good evidence that the effects don't exist. It's not proof though; they might just have not been very good experiments.
You give another example of choice architectures though I'm not sure if that's a nudge in the literature or not.
As an example to differentiate: drinking a homeopathic solution for health has no effect; driving a radium solution for health hurts.
I've seen the claim a lot but it all goes back to documents like this one which discusses strategies for communicating to increase compliance with lockdowns.
https://assets.publishing.service.gov.uk/government/uploads/...
Yes, the UK government received such shocking insights as "Messaging needs to emphasise and explain the duty to protect others", and "Messaging about actions need to be framed positively in terms of protecting oneself and the community, and increase confidence that they will be effective".
Of course, the government did pick and choose what to follow, so it would be absurd to say the entire COVID policy was "based on behavioural nudging". The UK's adherence to isolating after positive tests was thought to be one of the lowest of any country. When SPI-B pointed out that financial support would increase adherence to isolating, no reaction from the government. https://www.instituteforgovernment.org.uk/blog/government-su...
https://www.theguardian.com/commentisfree/2020/mar/13/why-is...
Or from the other side of the political spectrum:
https://www.telegraph.co.uk/politics/2022/01/28/grossly-unet...
The second one is about "deploying fear, shame and scapegoating" which the document I linked specifically calls out as a communication strategy with more downsides than any of the others they mentioned. However, Priti Patel just can't resist such activities.
I hope this doesn't lead to weakened immunity overall in the population. If you wear a mask every time you go out into the world, that doesn't give you much of a chance to build up acquired immunity to all the other bugs that are out there. There are stories from the early 1900s of native americans coming out of the woods and joining western society. They of course have spent decades in isolation versus just two years, but that's enough for them to end up perennially sick and in poor health when they were actually integrated into western society, and eventually die young of common disease. A lack of acquired immunity is what killed Ishi: https://en.wikipedia.org/wiki/Ishi
If I’m interpreting this correctly(and I by no means am sure that I am), I infer that they are saying in a fair publishing environment you’d expect to see more results that show less decisive results, therefor the current set of results is likely biased.
Couldn’t this bias also happen in the other direction? It sounds like they’re saying the results are too good and don’t match other scientific patterns of publishing results.
Whether publication bias is the explanation in this example, I don't pretend to know.
But imagine that a company would just as soon not pay out more 401k matching than it has to, so it makes the default zero. (Which of course is often the norm for different reasons.) That's as valid a nudge as anything but we shouldn't be surprised if a lot of people don't go with the default.
We probably also shouldn't be surprised if a lot of people maybe wouldn't go with a maxed out default.
Defaults wouldn't be nearly so powerful if they weren't typically chosen to be fairly reasonable for the average person in the target audience.
A funnel plot plots effect sizes on the x axis and precision on the y axis. The most precise studies should be tightly grouped around the meta-analytic average effect; the least precise studies should be spread more widely. This forms a triangular, funnel shape. If no publication bias exists, the spread of studies below the magnitude of the average effect should be comparable to the spread of studies above the magnitude of the average effect.
If there is publication bias, then the points that would form the left (without loss of generality; right if negative effect size) portion of the funnel will not be observable.
There are issues with funnel plots and there are other diagnostics but I hope this provides insight into one of the tools used. Notably, as a diagnostic, funnel plots work whether the true effect is positive, negative, or null; they assume only that the underlying assumptions of meta-analysis are true (that studies represent a sample of the same, true underlying effect -- other diagnostics and corrections exist when this is violated)
I'm not sure what theta is representing and only skimmed the paper, but especially in social scenarios and across social papers, seems unlikely to assume the same distributions and parameters across tasks & populations. Sometimes comes down to 'is there any effect??' and sometimes a precise notion of effect size in a lucky/clever specific scenario. Likewise, social science is one of the hardest fields to setup a good experiment, and few publications accepts negative results, so mostly only 'good' p-value ranges getting published seems normal. The Wikipedia page on funnel plots shows, afaict, the same criticism of the technique.
Whether about the effect size or how it is reported, funnel plots seem an inappropriate choice for debunking something as general as 'nudges' across heterogeneous studies. Skimming made the metaanalysis feel rather lazy (lack of cross validation, interpretation, ...). Not my field, but I would have had to do some digging before accepting this metaanalysis in review, and by default, would be 'not ready'.
I'd also add that this is a paper responding to an existing meta-analysis, so the claim that it's impossible to get a quantity of interest for a meta-analysis because the constituent studies are being inappropriately aggregated is itself an argument in favour of this paper's rejection of the original paper's finding.
I personally just rejected an m-a at a social science journal on thursday because I thought it suffered from unaddressed garbage-in garbage-out along the lines you mention (non-experimental data, no attempt to pin down causal identification, inadequate qualitative discussion of risks of publication bias, unclear QOI), so I am sympathetic to the criticisms, but just know that there are next steps. :)
it's another to reject the underlying phenomena because of metaanalysis.
In addition, they don't seem to have shown that the technique they're applying actually works for modelling the distributions that they're analyzing.
They succeed when lots of people say “what?! This broadly accepted idea is wrong?”
It’s the equivalent of a Buzzfeed headline, even if backed by thoughtful research. The new research may be correct in invalidating the prior experiment’s evidence but the reality is that we all know the “nudge” idea is useful at a practical level. If I ask my two-year old son if he wants milk, the odds that he is drinking milk 5 minutes later skyrocket. The same principle applies to people making all kinds of choices - from buying insurance to picking a University.
Choice architecture matters and we all know it.
So you're saying that we need to account for publication bias when reading a paper proposing a new way to evaluate publication bias?
And boy am I struggling - I am amazed it's even possible to group all of these studies under the same umbrella unless that is "misc".
Claiming that how people choose to treat their cancer, portion sizes at restaurants, rural Kenyan maternal health and Dutch childrens vegetable choices are even in the same field seems - incredible.
Maybe I am agreeing with the study in a roundabout way. If all of these things are under the heading "nudge" then it is too broad a heading. It's probably impossible to say one way or another that nudging works because you can never unpick all the confounding factors. Did the Dutch children have a popular TV show about vegetables while the Kenya media ran months long articles about unsanitary hospital conditions?
With my cynical hat on Nudge is a way for politicians to try something even when the real fix is intractable. I don't oppose "do something positive" - Injust oppose abusing power, violence for political gain, and all the other reasons why we can agree on a nudge in the right direction but cannot agree on a structural fix in the right direction.
I guess if they worked then they would solve the problems without structural chnage and so would defeat the forces that benefit from the status quo. so yeah. it does not work.
looks like we will have to go back to the old politics and revolt.
I'm one of the authors of the reply and it was very interesting reading so many diverse thoughts and comments. I would love to respond to all of them, but it would take ages. Luckily, Stuart Ritchie (@StuartJRitchie) wrote an awesome post on his substack (https://stuartritchie.substack.com/p/nudge-meta) that goes much deeper and adresses many questions and the fair critique raised here.
Also, note that there is only a limited amount of information and nuance you can comprise into a strict 500 words reply limit in PNAS, which is the reason we focus only on one aspect of the original meta-analysis -- publication bias.
Cheers, Frantisek
We should spread this widely in the hope that the pop ups and banners die off.
Or just that people want to believe it does?
> A newly proposed bias correction technique, robust Bayesian metaanalysis (RoBMA) (6), avoids an all-or-none debate over whether or not publication bias is “severe.”
Absence of evidence doesn't mean it's not true. It doesn't even imply it.
I tried looking but all I can get are articles citing the original Thaler claims/studies.
First, the term 'nudging' is a misnomer. Let's call it what it is - manipulation. Manipulating the options or defaults to some other set in order to achieve a better outcome for someone...
Well, who is that someone? The government?
Who says that their values align with mine? I wouldn't have responded as the government did to the pandemic, but their nudge units went into overdrive nudging people into vaccinations, etc. Is preventing access to bank accounts for protesting government actions (as in Canada) a 'nudge'?
Can I challenge the promoted values? If the state apparatus has its own values and agenda, how do I get to state mine - where is the values/ethics discussion being had, and how do I get my say? I find the promoted values Orwellian, communistic, overly progressive - one for all, but not all for one... is that opinion fair to hold? Or must I be nudged over the cliff?
Aren't we really just talking about soft-sell authoritarianism here? Weren't we just meant to vote for people, not have a perpetual nanny state guiding us?
What about all other public health campaigns? Drink driving, cancer screening, anti-smoking? Not everyone will want what they're pushing but we mostly agree that's a good thing for people, and that's why we let the government promote those things.
Think of how popular it became as a field in the last 50-100 years as the populace became less religious. The US adult population recently crossed a threshold where <50% believe in higher power now.
No science gives social scientists higher powers of forecasting human future, yet we took the ideas and applied them with the same conviction some believe in gods, in the same way; a network of randos spreading their gossip, wrapping it in technical jargon biased by past ignorance.
Consider how much of this work was being leveraged against an ignorant public with no opt out button, via print and TV. How is that informed consent?
Social media comes along, upends those forms of media, creates a new meta awareness we lived in a society policed by high minded but normal people. That awareness means we can opt out of being influenced by intentional nudges, same as we opt out of believing in intentional nudges to abide higher powers.
Social science “worked” when the masses were unaware it was happening to them. As the public has become more aware of how it works, it’s all Soylent green; just people.