133 karma · joined March 10, 2009
We'll look into making the chat less obtrusive. I also don't love aggressive chat boxes, but didn't realize Intercom did that on mobile.
I don't accept the characterization as a "brogrammer without ethics."
I don't think anything I say would change your opinions, but I'm willing to answer questions if you have any you want answered.
I'm Dan, CTO at Condorsay. We've built a tool that helps you make decisions, alone or with others.
* Summary
The short version of our tool:
- pick a goal (something to decide)
- pick factors important to the decision (helped by GPT-3)
- pick options (helped by GPT-3)
- use pairwise ranking to learn what you or a group think of the options
- see the results, including text notes (if you made the decision with others) and dissenters (people who disagree with the group)
NOTE: it does require a Google login to get past the "factors" screen. However, you get 5 decisions for free, so you can do everything for free. I read login makes HN cranky, but that is how our tool works. (A decision has to be owned by a user in the DB. Also, we are protecting against someone burning tons of money on GPT-3 calls.) Please don't be cranky.
* Motivation
We believe that decision-making could be improved by a tool that puts structure around it.
Some benefits:
- clarity: structure and record your decisions, alone or in a group
- focus: force hard choices with pairwise comparison
- revealed preferences: you actually don't always know what you think, but you learn through the simplest possible gut-level calls (pairwise choices). By the end, the results make sense, even if they aren't what you thought at first.
Further benefits if you’re making a decision in a group:
- alignment: get a group of people on the same page, by having the most important discussions quickly (i.e., where people disagree)
- independence: express your preferences before you see anyone else’s, to avoid information cascade
- asynchronicity: coordinate people in our increasingly remote-work world, mobile-friendly
Our scoring is straight-up Analytic Hierarchy Process (https://en.wikipedia.org/wiki/Analytic_hierarchy_process). For AHP in Python, we use the lovely https://github.com/PhilipGriffith/AHPy.
Happy to talk about paired-choice decision-making algorithms, AHP, GPT-3, or other TLAs (three-letter acronyms).
* Background
James, Andrew, and I worked on Barometer, an effort to use Facebook measurement to promote or demote ads to help defeat Trump in 2020. Even in that effort, we picked what to do next in various different ad-hoc ways. James thought, "Why isn't there a good tool for this?" and convinced us we should try to build it.
I look forward to your feedback and questions. Please don't be cranky.
Prisoner's dilemma problems are hard to solve.
"The University of Minnesota Department of Computer Science & Engineering takes this situation extremely seriously. We have immediately suspended this line of research."
I thought you meant split the infra pieces into different software components (process, machine, etc.), not split them by domain. Can you give an example, and why it's good to split by domain?
By infra, you mean for example a postgres DB? That makes sense. Don't want marketing to take down your website with a mail merge.
If your benchmark is plotting 2 million points, it may also be slow. However, I find its ability to rapidly prototype tons of different visualizations useful. I try to avoid needing to plot 2 million points anyway.
"Learn how looking for anomalies in your daily business metrics can help you more quickly find, learn about and respond to changes in your business."
2. To compute statistical power, you need an estimate of effect size. For many experiments, we don't know; we could run pilots, but in the naive case that doubles the number of experiments to run, and time to wait. By default, we recommend a sample size that will detect a certain effect size (but not smaller ones). That is, we decide we are not interested in small effects, because it makes our life simpler.
In the best case, it might be more robust to low-data outliers that cause random blips that look significant, as well as addressing the problem of looking at an experiment over and over (which frequentists are uncomfortable with: more chances to succeed).
However, it is non-trivial to understand how this works for us in practice. Examples: we want overall results as well as days-in results (to look for novelty effects); we would have to choose priors with consideration, because they have a huge impact on most results; etc.
This takes time and effort, and there are a lot of other things competing for those resources.
In each experiment, we assign users randomly based on a hash of experiment name and userid. So, we assume (and have tested) that the groups of one experiment have a fairly uniform mix of users from the groups of any other experiment. This means the primary effect of being in a particular experiment group G should be much larger than the effect of being in G and some other experiment's group. So, we usually analyze the effects of an experiment independent of any other experiments. Millions of users helps here.
I think this is how trials normally work: assume that random assignment into groups will spread all other factors equally across the groups.
There are experiments which depend on each other (change the same UI, one assumes a precondition of another), and that situation is complicated. We try to avoid it when we can, and negotiate carefully when we find ourselves in it.
You might also be referring to the fact that looking at hundreds of things will find "statistically significant" effects simply at random. As my advisor said, "95% confidence means wrong 1 in 20 times." That is a risk to be managed. We always have to ask things like: is this a reasonable outcome for this experiment? Do we have other corroborating evidence? Is the effect consistent? Do we have reason to disbelieve the result? Some fraction of our experiments are run incorrectly or have a broken implementation; we detect some, likely not all.
Also: is the risk higher to run an experiment that may be misinterpreted, or not to experiment at all and just release things? We like to learn from data if we think we can. There is art in the tradeoffs.
The problem is there is no reputation on craigslist, so people can be as bad as they want.
I can't compare that to ebay, tho. Maybe ebay is worse for a seller.