HNHacker News
TopNewBestAskShowJobs

ddp26

762 karma · joined April 10, 2012

Dan Schwarz. Co-founder of futuresearch.ai. Previously Metaculus, Google.
submissionscomments
ddp26··on Gemini 4 Argon
Isn't the important takeaway here that Gemini 4 is not released and has no planned release date?

This is marketing from Google, not a competitive offering

ddp26··on Artificial intelligence now beats some of the best human forecasters
Yes, I heard from one first-rate forecaster that he thinks AI forecasters are especially weak in predicting big disruptive changes to the world.

Hard to study this, obviously!

ddp26··on Artificial intelligence now beats some of the best human forecasters
Right. What's really surprising is how much better the best are. Human superforecasters, and prediction markets are surprisingly accurate too.

We could live in a world where things are much more chaotic, and the best humans (or AIs) would only be slightly better than chance. Evidently the world we live in is pretty darn predictable.

ddp26··on Artificial intelligence now beats some of the best human forecasters
As someone who started working on AI forecasting 3 years ago, I can confidently say that most people did not expect AI to beat Tetlock's superforecasters, Metaculus pros, or prediction markets as quickly as it did.
ddp26··on Artificial intelligence now beats some of the best human forecasters
The Economist has actually published other human forecasts many times, e.g. Metaculus or Good Judgment forecasts. They do year-end forecasts too.

Whether they draw on AI or other humans seems immaterial to the quality of their reporting.

ddp26··on Artificial intelligence now beats some of the best human forecasters
I know this is tongue-in-cheek, but I think your idea could actually work, but not in financial markets. (The "keynesian beauty contest" of trying to predict what others think been played out to death there.)

You could train a model to anticipating scientific trends. Or policy trends. Others will definitely use mainline LLMs to make decisions there, so they may be more predictable now!

ddp26··on I can't stop thinking about Papua New Guinea
I spent a summer in PNG in 2010, on the island of Karkar. It was wild.

One of my most formative memories was finding out that few of the people who live there, even those who are literate, had books and some DVDs, were high school students, even had occasional access to computers with internet, knew that humanity had landed on the moon!

ddp26··on Gemini 3.8 Flash and 3.8 Flash Cyber
There must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model.

What's the simplest explanation?

ddp26··on How accurate have Ed Zitron's AI skeptic predictions been?
Even though these predictions turned out mostly wrong, we should not castigate people for publicly forecasting! That is virtuous, and more people should do it.

Thank you Ed!

ddp26··on Gemini 3.7 Flash
What are we to infer from no release of gemini-3.5-pro, but frequent releases of smaller flash models (presumably from the same large pre-training run?)
ddp26··on Goodhart's Law Comes for Every Benchmark You Trust
Not forecasting though. You can't goodhart predicting real-world events
ddp26··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
This seems bad for AI safety/risk. Does DeepMind have any checks on model alignment now? What's stopping them from using AI for military/surveillance purposes?
ddp26··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
Would it though? Meta and Microsoft have had very scandalous AI things happen, and their shares didn't tank (or quickly recovered)
ddp26··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
I would have thought Jeff Dean would never ever leave Google. What on earth is going on?
ddp26··on Data centers have hiked electricity prices on the public by $23B
Isn't this the same as saying "utility regulators delaying connecting new power to the grid hiked electricity prices on the public by $23B?"

When my apples are expensive, I don't generally grumble about all the demand from pie makers. If they demand more apples, new suppliers should come in to restore the price, right?

ddp26··on The Tower Keeps Rising
> I can ask an agent to add OAuth, you can ask one to add caching, and somebody else can ask one to rebuild the database from first principles and make the UI pink. Each change can be reasonable in isolation.

But this is just bad vibecoding? This would be bad if humans did it too. With agents or humans, you need to coordinate.

ddp26··on Proof of care in the age of AI
When I know something is (primarily) AI generated, I lose interest.

The exception is when it's about a niche I care about, e.g. an analysis of opening trends of early world chess champions. I'll read AI on that for an hour.

My sense is that, for most writing, it's fundamentally interpersonal, the information is about the author as much as it is about the world.

Maybe this flood of slop will cause people to care more about the substance of the writing, not the perspectives of the writing.

ddp26··on GPT-5.6
Is it possible GPT-5.6 is not a very aligned model?
ddp26··on GLM 5.2 and the coming AI margin collapse
People have been making claims about the commoditization of llms since chatGPT, and they've been wrong every time as quality and prices and differentiation have increased.
ddp26··on The AI Superforecasters Are Here
But Scott's point is more: why even have markets? Once you have the superforecasting available on the questions you care about, why do you need to publish it for everyone to also react to?
ddp26··on The AI Superforecasters Are Here
Almost by definition, once AI forecasters are in the market, they won't (all) be beating the market.

But why evaluate AI forecasters by beating the market? Do we evaluate deep learning by whether hedge funds make money from it in the markets? These things have far, far more utility outside of finance.

ddp26··on The AI Superforecasters Are Here
Doesn't this argument prove too much? Why does AlphaSense sell their company research instead of using it to trade themselves? Why do people work on open source time series forecasting packages instead of quietly using them to trade?
ddp26··on The AI Superforecasters Are Here
Indeed they are! It's funny, in Sept 2024 I and others wrote about how the AI Superforecasters _weren't_ here, despite several claims that they were: https://www.lesswrong.com/posts/uGkRcHqatmPkvpGLq/contra-pap...
ddp26··on Podman v6.0.0
Yeah, a great developer I know showed me how he could use it to get a safe dev container for Claude Code, in a way that wasn't doable with Docker.
ddp26··on Previewing GPT‑5.6 Sol: a next-generation model
Fair point.
ddp26··on Previewing GPT‑5.6 Sol: a next-generation model
Is this the trend? There have been various points where one of Anthropic or OpenAI was substantially ahead. Sure, many times they're close, but now doesn't seem like one of them.
ddp26··on Previewing GPT‑5.6 Sol: a next-generation model
Based on my conjecture that Anthropic is ahead on AI research, and that OpenAI doesn't know how to make Fable-class models.
ddp26··on Previewing GPT‑5.6 Sol: a next-generation model
I'm going to pre-register my prediction that GPT-5.6 Sol is significantly behind Claude Fable 5, as evaluated by general consensus once time has passed for people to get familiar with both.
ddp26··on World-Modeling the US vs. Anthropic on Claude Fable
If it's helpful, I'm still holding at July 9 as my median date that Fable gets re-released to Americans, the news of the last 24 hours didn't update the model meaningfully.
ddp26··on Mark Zuckerberg directed Meta to create a prediction markets app
People forget that Meta already did this years ago, before prediction markets became the next big consumer trend for them to chase.

The app was called Forecast, and launched in June 2020. (Around the same time that Kalshi and Polymarket launched, actually!) It was framed as a way to make the comments and activity on Facebook actually productive rather than toxic, and build expert reputation signaling mechanisms.

I think they sunset it after about a year.

Page 1 of 6Next →