Very cool
128 karma · joined January 13, 2026
Very cool
For example, that french dude that bet on a prediction market what the temperature would be, then broke into the weather station at the airport with a hair dryer. The prediction was still specific and falsifable!
One pony’s trash is another pony’s treasure, so I guess here one persons pedanticism is another persons hobby.
Bad (form) prediction: OAI is gonna be wobbly in a bit Good (form) prediction (could be totally wrong): OAI as we know it today is going to collapse due to running out of money around/at 20xx.
And sure we can be pedantic about detail, but the litmus test is: is the prediction useful if you had a crystal ball and you could know if it was true/false a priori?
This is a spectrum of course: a prediction that OAI will collapse is probably right as _eventually_ all companies come to an end, but under that interpretation, the prediction is useless. It's more signal / useful / falsifable to say OAI is going to collapse around/at <year> due to <thesis>.
Anyway thats my two cents. Its fine to outline forces and trends, but when you make predictions there are useful (better, falsifiable) and useless.
But predictions need to be specific and falsifable. If not, its just rag-chewing over a beer (luv that shit, but i aint predicting on taco tuesday). If they arent falsifable, then its not a prediction.
I think a lot of the HN comments generally can be described as one camp which cares about and enforces the rigor of predictions and trying to direct limited ear-time to voices which tend to get predictions right, vs the other camp that puts more weight towards directional accuracy.
On the overly-literal / narrow thing, I think thats the culture around evaluating predictions overall. Like, all those posts around christmas where people make predictions and evaluate how last year went. The rigor is the norm.
(you either go the path of a burner or a birdie, I don't make the rules)
You can see the implementation here: https://gitlab.com/cryptsetup/cryptsetup/-/blob/main/lib/cry...
Iiuc youre saying: its more cost-efficient to waste water using evaporative cooling so thats what we'll get, not that a closed loop with a heat exchanger is technically infeasible?
Also whats your cutoff for 'acceptable' speed? I would have said 25tok/s.
Most of the benchmark improvements afaict are in agentic and instruction following benchmarks.
Perf improvements seem to all come from training?
But please keep writing, I know its super hard to put yourself out there and make content!
1. The specific bug isnt mentioned 2. (If youre game) a model with a knowledge-cutoff date before the report is used
I couldn't find it, so its unclear if the prompt was completely "make a test suite" or was lead towards finding it in the first place, which wouldn't be a fair test.
The closest I mention of the prompt I could find was:
> Then I asked it to write a simple workload which exercised the WAL insert and checkpoint code. Notably, this is a completely generic workload.
With a skeptical lens, unclear.
Given its basically an O(1) lookup with disk space being the main constraint, I was curious if you've tried ablating engram layers and sizes across your setup?
Also, why mHC over attention residuals?