HNHacker News
TopNewBestAskShowJobs

jellyberg

45 karma · joined June 13, 2023

Building fatebook.io and quantifiedintuitions.org
submissionscomments
jellyberg··on A ragtag band of internet friends became the best at forecasting world events
Forecasting tournaments: https://www.metaculus.com/ https://www.infer-pub.com/

Forecasting your own life: https://fatebook.io/

Training your calibration: https://programs.clearerthinking.org/calibrate_your_judgment... https://www.quantifiedintuitions.org/calibration

Books to read: Superforecasting, Scout Mindset

jellyberg··on A ragtag band of internet friends became the best at forecasting world events
See https://samotsvety.org/track-record/
jellyberg··on Compare how Gemini Pro and GPT 3.5 answer the same questions
Ah thank you, I missed this!
jellyberg··on Compare how Gemini Pro and GPT 3.5 answer the same questions
There are four examples in the first category (imagination), there are five more categories with lots of other examples.

Currently, there's no API for Gemini Pro, so you need to use bard.google.com to ask it questions.

jellyberg··on Compare how GPT-2, 3, 3.5 and 4 answer the same questions
Yeah - worth noting that we use temperature=0 for reproducibility while ChatGPT I think uses t=0.7. We also prefix the prompt with few-shot examples of questions and answers with chain of thought examples to elicit the models' full capabilities.
jellyberg··on Compare how GPT-2, 3, 3.5 and 4 answer the same questions
Yeah it's hard to compare across models, interested in suggestions here.

We give all models a bunch of few-shot examples, which improves GPT-3 (davinci)'s question answering substantially. GPT-2 sometimes generates something that answers the question, sometimes it's just confused. Click "See full prompt" to see the few-shot examples that the models get.

Our goal was to exercise the full capabilities of each model.

jellyberg··on Show HN: Fatebook – Superforecasting in your Slack
Some examples of different kinds of resolution criteria for forecasting questions:

Personal mechanisms

“If we choose this HR provider, will I think it was a good idea in two month’s time?”

“Each day I’ll write down whether I want to leave or stay in my job. After 2 months, will I have chosen ‘leave’ on >30 days?”

“Will volunteering abroad make me all-things-considered happier?”

“Will I judge that AI was a major topic of debate in the US election?”

Objective proxies

“Will our user satisfaction rating exceed 8.0/10?”

“Will Our World in Data report that >5% of global deaths are due to air pollution by 2030?”

“Will I still be discussing my fear of flying with my therapist in 2024?”

Experts

Consider only actually asking an expert on ¼ of the questions you generate, chosen randomly, to keep your accuracy incentive and get feedback on your predictions, but minimise the cost of consulting experts or mentors.

“Will my mentor agree that pivoting now was the right choice?”

“Will the rest of the team prefer this redesign to the current layout?”

“Will anyone on the animal advocacy forum share evidence that convinces me that abolitionist protests are net-beneficial for the movement?”

“On December 1st, will Marco, Dawn, and Tina all agree that the biosecurity bill passed without amendments that removed its teeth?”

Crowds

“Will >80% of my Twitter followers agree that I should keep the beard?”

“Will a Manifold market on this question trade at >90% YES in two weeks?”

“Will social media posts about our product be more positive than negative in the first month after launch?”

“If I survey 40 random Americans through Mechanical Turk, will our current favourite name be the most popular?”