Wizard for Mac – new kind of statistics program
wizardmac.com
wizardmac.com
SPSS is a classic example - and the social sciences have had a series of discredited papers over recent years due to poor application of statistical methods. Just because you can put a dataset in and get values out - even values that look significant, it doesn’t mean they are - did all of the assumptions and requirements of the methods/tests you used hold true/pass with your data?
So, while I haven’t looked closely at this tool, all I saw was talk of the interface and the ease of getting results even if “you don’t know where to start”. That scares me. Especially when you start talking about applications in domains like medicine. That could be lives in the balance. Would we want civil engineers using tools like this to build bridges? Or people designing nuclear reactors?
To me, this is worse, not better unless it somehow helps you actually understand when, why and how to correctly use these tools.
Promoting ease of use of tools and simplification of harder problems is great, but this is a really, really dangerous thing to make easy and oversimplified.
Nope nope nope nope nope.
SPSS has, shall we say, a less than savory reputation.
All this to say that there is a market for something much friendlier than R. R is used by pure statisticians, data scientists and the like, but most social scientists prefer Stata, which has pretty legit statistical routines as well as a point-and-click UI.
I think it surprisingly makes a lot of sense to social scientists, in spite of seeming backwards to computer scientists. I remember being in school and using STATA, and just plugging in numbers to get through exercises too quickly to bother understanding what the labs were about.
R seems to make you stop to realize what you are doing at the right moments, then uses a lot of magic to abstract away almost everything else.
R is objectively a bad programming language. However, it is by no means inaccessible. I have no statistics background whatsoever, and I managed to learn enough R to be dangerous in a mere week. Other than the 1-based indexing and the utterly disgusting dynamic dispatch mechanism (you could simply not to use the latter), R is surprisingly pleasant to use. What I enjoyed the most is that vectors and matrices are first-class values, not objects that are referred to through pointers. It's probably copy-on-write under the hood, but I don't need to care. Hallelujah!
As you note, getting an answer and getting the statistically correct answer aren't the same thing.
You don't want to design the control room for a nuclear reactor like you design a webpage, because you can under no circumstances sacrifice correctness for ease. A lot of times, it's about making the wrong thing really hard to do, with multiple layers of fail-safes. The potential gain in productivity when designing for ease is vastly overshadowed by the risks associated with making an error.
When we're not talking about nuclear reactors, but—for example—statistics, where consequences are abstract, it might get tempting to err on the side of ease. In fact, any domain you don't understand well enough well tempt you to design for ease over correctness.
Ideally, the designer should understand the domain at hand well enough to create a design that makes it easy to be correct, and hard to make an error.
There's a lot of stuff out there that has been neither designed for ease nor correctness, however, but which has a design arbitrarily dictated by the table layout of a database, or some other random technical constrain that has nothing to do with the problem domain.
Just because someone has spent a lot of time learning how to use a difficult tool doesn‘t mean they understand what they are doing any better.
What is wrong with having a fast and easy to use tool which makes data analysis accessible? If people misuse their data or their tools, that's their fault.
All of those things and more exist for the protection of the users and/or general public - but many people hold the attitude that it should be up to the user to do or not so it right.
We have a history of banning or socially stigmatizing dangerous products and practices - see 3 wheel ATVs and lawn darts, for easy examples.
Similarly, and less extreme, standards and best practices exist exactly because it’s often pragmatically dangerous to expect the end user of a product to be an expert enough to avoid hidden or infrequent but high risk dangers, much less very common and high risk dangers. (See examples everywhere from medicine to aviation or the national elictrical code).
Making it easier to shoot yourself in the foot isn’t a virtue on its own...
The difficulty of learning the software doesn't says much about the quality of the ideas, of the research design, of data gathering and interpretation.
Like it or not, SPSS made science much easier. And we should work in that direction. Not creating more complex tools for no reason. A software can and should also be a pedagogic tool, guiding people through...
No - it didn’t. Science isn’t running stats tools. It’s knowing what tools to use and using them properly so that you get valid answers.
If you haven’t been paying attention, the social sciences have been plagued by an epidemic of failures of reproducibility - see: https://www.nature.com/news/over-half-of-psychology-studies-... for reference. Only 39% of studies were found to be replicable and much of the blame falls to poor methods and techniques like p-hacking.
> And we should work in that direction.
Yes, we should - but dumbing down interfaces in ways that don’t actually provide good guardrails/handholding to help with proper application is not making good science easier.
> Not creating more complex tools for no reason.
Nowhere have I or others suggested that. But a tool that makes it simple to get an answer is not the same as one that makes it simple to get a correct and valid answer.
Science is a means to an end in many cases. Medicine, for example, is about better human lives. Shoddy research methods could do anything from giving false hope to actually endangering lives.
Everyone I show this software to goes: "Whoa, what is this?" I would recommend checking it out before dismissing it.
I’m not going to write a script or configure complex software for a quick check. Wizard is perfect for that. It’s also super fast at scanning through really enormous CSV files, then generating plots for every parameter.
Every coworker who sees me using it for these tasks wants to know what it is, or how I send out relevant plots so quickly when new datasets are available. It’s really a fantastic tool for quick work.
There seems to be a lot of negativity in this comments section about misuse of statistics. I think people are missing the point. Easy tools make data analysis more accessible, but misuse of data is the fault of the user, not the tool.
Forgive me if I seem overly aggressive but I have grown weary of my and my colleagues' profession being side-lined and belittled. Politicians, administrators, and even some other scientists see statistics as merely a badge to be placed atop their own work for validation. Well, ladies and gentlemen, statistics is more than that. It is an empirical science in its own right.
Of course, it doesn't always take a statistician to do the necessary statistical work. I am no physicist yet I can certainly apply the Clausius Clapeyron equation as needed. Likewise, I expect many (perhaps most) scientists to be able to apply an ANOVA or simple regression as the need arises.
HOWEVER, the lack of intellectual humility on the part of so many non-statisticians when applying statistical tools to their own work is maddening.However, if we consider judging whether a statistical method will be useful for the world as a part of statistics, then that part sometimes is an empirical science.
That is, the statement "I should use X technique because it performs Y% well" is sometimes an an empirical statement.
There is nothing about easy-to-use that precludes understanding, and certainly nothing about difficult-to-use that promotes it. They are largely orthogonal. Using R doesn't make you a statistician, any more than using C++ makes you a software engineer. If anything, a simple interface can reduce the number of ways you can shoot yourself, and leaves more time to focus on the problem.
Being easy-to-use may be the difference between some analysis and no analysis, or at best, analysis by spreadsheet.
And finally, this is Hacker News. The author wrote this software and makes some money off it. Great. Isn't that what this place is all about?
This is much simpler. You just input the data and it gives you charts right away.
It isn't R, SPSS or Minitab. It's brilliant at what it does and I love it. I've been using it for 3 / 4 years and wouldn't swap it for any other tool.
Making good-looking figures nowadays is not a selling point anymore. If one's willing to script rather than clicking-the-mouse, Prism, Igor Pro, Origin Pro, Matlab (pricing from low to high) all can produce great figures and solid statistical test for people out of the statistics field. But nothing these days beats R for versatile of statistical tests.
[1] http://www.evanmiller.org/ [2] https://www.youtube.com/watch?v=TzJMFxj7GRI
Maybe this makes the trial versions better?
-Firstly, how does the visualization know the respective population or sample size from which the summary statistics and intervals are to be drawn?
- The demo used a pie chart to try to display summary stats and confidence intervals from the general social survey. Aside from professional statisticians general dislike of pie charts, you cannot plot confidence intervals in this way into a pie chart just by inserting 'more white space between the slices'. There's only 100% of the area of the circle that you've got to play with, so any attempt to increase the 'white space' between the slices necessarily warps the real estate remaining to represent each actual slice.
- Honestly, I see this tool likely to be used by people who participate in the practice of p-hacking, whether deliberately or not. The ability to throw lots of simple models quickly at lots of data mindlessly reporting some notion of statistical significance is dangerous. I'm assuming your stats are not (cannot) be adjusted in any fashion to implicate what you're really doing by using an automated model-building/reporting regime in this way (potentially running heaps of models on heaps of data until you find one that appears 'significant' based on a statistical test designed under the assumption that this is NOT what you're doing). No where did i see any application of train/test, sample/resample type methods to try to control for over-fitting in the prediction application or truly estimate how predictive/replicable such a technique would be in the real world.
While I appreciate the work done required to put something like this together (a lot of it looks like a gui interface to my own exploratory functions/scripts in R, for example), i genuinely believe this approach is more dangerous/likely to lead to false conclusions than helpful.
As I said elsewhere, a simple interface doesn’t necessarily mean a correct outcome - In short, most stats software is solving the wrong problem. Many (especially this one) make it easy to get an answer, right or wrong. I’d rather see them make it hard to get a wrong answer - or, perhaps, hard to get an answer when you’re using it wrong.
Honestly, I think this is a tricky human problem, not a tech problem.
My reasoning is thus. I've been known to argue that even R is bad from this perspective: I view its success (apart from its free/OS nature) due to the fact that it carries with it libraries, a functional flavour, tied around a core engine/philosophy of implicit actions preference to result production rather than bothering the user.
These enable a person (with just enough knowledge to be dangerous) to load an externally authored package (that has neither been tested nor verified) with one line, load a dataset with one line (which silently corrupted or changed something during import), and apply a function in one line (which silently coerced objects/values in the background and unreliably expressed/suppressed errors and warnings). Where it does express warnings/errors, it does so unreliably/unhelpfully, so amateurs are led down the path of excessive warnings/errors where things continue on regardless: ignore them.
To many, what I've written up there is not a bad thing: you get statistical analysis in three simple lines for free, all while copying and pasting scripts from stack overflow or the internet.
Now, i'm actually working on my own hobby project of designing a language/library for my own use which is designed around fixing those principals: still function based, interactive, fast, no implicit coercion, allowing flexibility while imposing restrains and guarantees.
But I'm under no illusion that it would necessarily be popular if i ever realised it to the public. The attraction of R is that you get a model out in three lines, rather than 14 errors and no result telling you that there are issues involving realms of thought you didn't know you were ignorant in and you'll have to go away and study before you continue, or even that your data might not be suitable for what you're doing. It might not be a good piece of work, but cynically, for the type of person buying into such a mind-set of quick analysis via "darts thrown at an analytical wall", i'm not convinced that quality genuinely matters to them (even though it matters in the effects it has on the public further down the line).
And i've not even gotten into the problem that in the real world 90% of work is not analysis/modelling but in data cleaning/munging, critical thinking and technical problem solving, and that there's another layer of problems below that one which most academics and professionals rarely engage in, which is questioning the systematic/contextual nature of the data before it got to your data set (which no software/interface i'm aware of currently addresses and most people just blithely ignore).
Why not have the video at the top. Maybe pop right on my face. Actually, these are moments when I'd not mind a popup that takes focus out of a page.
Often, I have questions that take the format of "for each of the following categories please rank them between strongly disagree and Strongly agree". The way these questions end up in the SPSS files are typically as different variables for each row, and a 1,2,3,4, or 5 as the measurement along with the labels.
Frequently, I want to pivot those types of questions by a variable like Region or Number of Employees (categorical), and then see the resultant tables. This is never fun, and inevitably takes a lot of time.
As others have said, statistics is a careful business that doesn't necessarily warrant ease of access to all mathematical functions BUT, handling what SPSS calls "Multiple Response Sets" better would be a godsend, just for the data prep and visualization step. I still ultimately fall back on recoding these or leveraging the MRS functional in SPSS to get this done (sometimes this is better than just using pivot tables in excel).
It would be great to be able to specify this kind of thing in this program, since without it, you can't really use/trust the computed percentages in some question configurations. Take a real look at the SPSS Tables feature, the Multiple Response Sets, and then visualization of them, and consider how that data is actually coded in SPSS files (the common export of survey tools) and maybe you can improve on that feature (It shouldn't be hard, MRS is a pretty bad setup, but it gets the job done).
Some of the graphics need to be improved. It would also be great to see why this is better than tradition BI tools (Tableau etc) and what your unique value proposition is.
Wow, that's a great way to kill off any sympathy for him. Especially given that he's apparently had 4-5 years to dig himself out of that hole.
Along with Python and Julia.