6,610 karma · joined January 17, 2014
Some fun stuff:
https://borgcloud.org/speech-to-text at $0.06/h
Roxy: iOS hands-free voice AI: https://itunes.apple.com/app/id6737482921?mt=8
Turing Test Battle Royale: https://trashtalk.borg.games
meet.hn/city/43.6534817,-79.3839347/Toronto
Socials: - linkedin.com/in/victor-msu - reddit.com/user/lostmsu - github.com/lostmsu
Interests: AI/ML, Gaming, Networking, Programming, Research, Science, Startups, Technology
---
I even hinted at the meaning by using "spirits". Next time you may want to use AI to explain usage of words based on their context before jumping to telling me what to do.
I mean you don't even understand that spirit death in the context is not synonymous to bodily death, but a mere consequence of it. Really bad.
Same, but actually much of your entire life recorded and trained on.
Having "Wait, bar is not true, so that won't work" is not necessarily a correction. In fact, the problem is: across a long text it is a correction of a single mistake, but we are talking about thousands here.
But yes, of course that was a rough estimate. But the problem is - we don't really know what we are measuring here. Maybe there's a 2,000,000x difference of intelligence between coding indexes 52 and 50. By some measure that just feels small because that's how we process it akin to audio db.
Regardless the point is KLD and whatever they came up with is not meaningful. And they did not publish comparisons on real benchmarks.
KLD of 1%, or similar error metric that multiplies, on 10000 tokens would give accumulated error of 2,000,000%
What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?
The problem is that the pay for that goes to Amazon instead of your pockets. I think Brave was actually on to something.
Let's say it's variability (number of distinct if you prefer discrete) of stable responses to external stimuli.
You can't really know that either.
Just wanted to say that this is a very important point that I totally agree with. People are obsessed with KL divergence, but it is yet to be demonstrated to be a descent proxy for agentic coding benchmarks.
Still trying to lock in, huh?
Why is Qwen3.5 2B not in the table?
That's exactly the point. We know short context knowledge stuff does not regress with quantization. But I expect agentic intelligence to suffer greatly.
If I were to pick one bench, I would like to compare quants on TerminalBench Hard. But then Glimmer already loses to 3.6 27B on it by a large margin.