Introducing Guesstimate, a Spreadsheet for Things That Aren’t Certain
medium.com
medium.com
It will be a bit of a challenge to make it intuitive, but I'm sure I'll figure out something.
Just so you know, right now there is a second distribution, uniform. You can access it by clicking the 'normal' icon in a normal distribution, where a short list is shown below.
* Prediction markets are usually for binary outcomes. I imagine the most useful role of binary variables in Guesstimate would be to mix two different distributions. "If Clinton wins, student debt in 2018 will look like distribution A; if Sanders wins, student debt in 2018 will look like distribution B".
* I'm not sure how Augur (or any other market) reports likelihoods, but it's good to keep in mind that market prices do NOT generally reflect any sort of average belief. See https://www.aeaweb.org/assa/2006/0106_1015_0703.pdf.
In the future, one idea would be to keep track of people's metric estimates in Guesstimate, and later score and rank them on how well they do. So if Charles always reports a 90% confidence interval that's far too optimistic, we could help adjust it automatically next time. This would also allow us to aggregate different opinions directly, essentially being like a mini prediction challenge. This would be a ways off though, and it really depends on what direction the product goes.
Only request would be to allow for private spreadsheets. I can download and run the code locally but this would help many people who are less tech savvy.
Great product - looking forward to seeing how it evolves!
Cool idea. Found this from a @worrydream tweet. Some comments after playing with it for a few minutes:
* A bit disappointing that I can only have uniform or Gaussian distributions. At minimum I'd like a binary distribution (coin flip, probably biased coin flip). A lot of things I would want to model need this (e.g., will we close this sale? will we close this investor? that kind of thing.)
* I'm really confused by the arbitrary two-letter codes assigned to things for formulas. Makes the formulas impossible to read. Why not just use the names I give to the cells, or something derived from the names?
Really nice start though! I'm co-founder & CEO of fieldbook.com, another spreadsheet-like tool, so I love information tools and anything that expands the mind's capacity. Best of luck and let me know how I can help!
In response to your points:
- Other distribution types are the #1 most requested feature at this time. Binary distributions are possible using functions instead of the built in distributions, though obviously not intuitive. It's built using math.js, which has several random functions of different kinds. (In this case, you could use '=randomInt(0,1)' to produce a coin flip.
http://getguesstimate.com/models/365
- The arbitrary two-letter codes were simply the easiest thing to begin with. I started with something derived from the names, but this presented problems with cells with empty names, especially ones that started empty and later become non-empty. Excel has a pretty sophisticated model for referencing cells. For the sake of getting something shipped, I started with a very simple one. Definitely an area to improve.
- I'd love to talk sometime. I'm also in SF, will send you an email. Thanks for the advice!
Re the two-letter codes, I totally get doing the simple thing just to launch and get it out there. We face a similar problem in Fieldbook, by the way. Would be happy to explain how our solution works sometime.
Give me beta access! I want to play with this thing.
Wow! The developer tools are amazing too: https://www.youtube.com/watch?v=stwlaJGeLoM
Of special interest are non-continuous distributions. How often have normal distribution reasoning failed in finance? Put another way, a user should be able to model a distribution himself.
When I built this, my first goal was to make any distribution run quickly. At this point I believe adding other distribution types will be quite doable, expect them shortly.
http://research.microsoft.com/pubs/208236/asplos077-bornholt...
I suggest trying it out. If nothing else, you may be able to begin with very simple models of the most important variables.
Feel free to play with it. You can edit it, just not save it. (I recently realized this was not obvious to most people)
As per the paper , you can choose arbitrary distributions , construct a fluent graph , run Monte Carlo simulation and get the result - |via http://bit.ly/hnbuzz01 |
Perhaps that field can provide a potential source of new names, when you decide to market this as a company.
I can't make guarantees about the distant future. There's a ton of work I would love to see happen with Guesstimate, and my guess is that much of it would only be possible if it becomes a company. This can still mean that it can be mostly open source, but I really have little idea what the situation would be at that time.
How does one tell guesstimate that there's a hard lower bound on a quantity. ie. Video Length is at least 0, because negative watch times are unphysical? I know the specified distribution in this case is very narrow (the video lasting between -1 and 0 minutes has probability ~0.000032). But the answer does come out to be 26±32, which includes a substantial unphysical region.
And, if I give a hard lower bound on Video Length, can it propagate that knowledge into an asymmetric error on Total time?
Right now the main distribution types are normal and uniform. In the video, I showed normal distributions, which have long tails in both directions.
In this case, a normal distribution isn't really correct, because, as you noted, being less than 0 is exceedingly unlikely.
I believe the correct way to deal with this is to use a lognormal distribution or something that has 0 chance of being less than 0. I don't yet have a simple way of doing this, but it's definitely on the agenda.
Just something to keep in mind when abstracting what a distribution is.
edit: though as a short-hand for entry, Gaussian is usually a pretty good guess. Is there support for µ±σ instead of [low,high] in the works? Or support for numerical distributions?
"Or support for numerical distributions?" - By numerical distributions do you mean discreet distributions: like, a 40% of being '8' and a 60% chance of being '6'? If so, the answer is no. However, if you use the ternary operator it is possible to do very simple versions of this now. We do support totally random picks of different numbers though, using the pickRandom([3,5,3]) function. http://mathjs.org/docs/reference/functions/pickRandom.html
Video: https://www.youtube.com/watch?v=w4fHGTsZZD8 Book: http://www.amazon.com/How-Measure-Anything-Intangibles-Busin...
http://www.theatlantic.com/politics/archive/2014/03/rumsfeld...
Those are standard epistemological distinctions, known (and written about) since at least the times of Aristotle.
[1] (I mean the philosophical essence of what he said -- not that he didn't tried to use it an an excuse for BS).
A few weeks back I used the tool to help a friend decide which mortgage option to take for his house. One house has a slightly lower APR than the other, but was a had a higher assistance fee.
After fiddling with it, it looked like the one with the lower assistance fee was the better option. But perhaps more important, it didn't seem like it made a big difference; perhaps around $200 after 10 years. This was a good indication that the choice didn't really matter; that it wasn't something to spend over a few hours worrying about.
I found, remarkably to me at the time, that at a single particular loan life, the APRs of a single lender all converged very accurately to a single interest rate. This told me two things: 1. Choosing a realistic loan life was key. 2. Many of the choices (points + fees vs. interest rate), were for most people illusory offerings to give the illusion of choice.
Surely, such a platform would make building an app 100 times easier. Not that building apps is a good use of our resources.
You can only access index N of a list whose type expression proves that it has at least 5/items. You can only access the fields of objects that are proven not to be null. You can only give as input to a function that expects an odd number a value whose type expression proves it to be an odd number. Basically, you can never have runtime exceptions.
If you don't like making types explicit, you could use implicit typing like in Haskell, or having values carry their types like in Ruby.
You might consider upping the run count, or maybe narrowing your bins for the visualization. Either way, it's great to see more tools embracing probability and uncertainty like this.
I think that this works fine for small models, which is much of what exists now. As there are larger models, I'd eventually like to offload calculations to AWS Lambda or something similar, so we can do far more.
image link https://camo.githubusercontent.com/8fd97a97fa656a1eb92294f0f...
It's definitely not mobile compatible yet. I think making it viewable on mobile devices soon is doable. Making them easily editable will be much harder / require a different interface.
And to comment on the project, incredibly interesting. In my experience non-quantitative people (non STE-grads) tend not to be able to, at all, model decisions, decision trees and event trees as distributions or probabilities. This kind of decision analysis can be sold at very high margins, because it's value can be very high for the right audience. Intuitively, if you can form a solid business team and do enterprise/policy/strategy level sales, a company around this has a market (none of the established competing software pointed out by others is very cheap, and none of it is very modern, accessible or intuitive). Another option, more suitable for a freelancer mindset (and with a wider distribution of non-bust outcomes as well as arguably higher EV...), is starting and maintaining this as an OSS project, hopefully with a wide and growing contributor base; and establishing a career as a consultant in the modelling/predictions space with Guesstimate as your resume and business card. This of course depends as well on your personal interest - software developer vs. data scientist vs. fledgling capitalist.
Also check out:
http://probcomp.csail.mit.edu/bayesdb/
https://github.com/taschini/pyinterval http://mavrinac.com/index.cgi?page=fuzzpy
Also, use pandas ;-)