1,729 karma · joined September 3, 2022
[0] https://www.europeanpaymentscouncil.eu/what-we-do/sepa-direc...
If you're running Qwen3.8-27B on energy-efficient hardware like a Mac or a DGX Spark instead of an API (likely running on H100s), I'm sure you're having much more of an impact than you would by switching to, say, a 9B coding-only model on the same hardware. The thing is, I think you won't be able to go orders of magnitude smaller, because a lot of the usefulness of LLMs comes from emergent smartness, and you typically need a minimum amount of complexity to see such emergent phenomena (and I think we're pretty far from understanding this kind of emergence, much further than from the next model generation that annihilates the current one on benchmarks yet again).
No. It means that the one model has 2.4 trillion parameters while the other has only 27 billion. I don't know the details about their architecture or training, but presumably they used the same or similar training sets for both and a conceptually similar architecture, scaled down. I'd guess they also have some techniques to re-use some of the work done for the big model for the smaller versions (if anyone knows more about this I'd be interested). The architectures cannot be identical by definition because then the parameter count would be the same. Subsetting the data to such narrow fields as you describe could risk losing some edge, there are a lot of emergent capabilities in those models and I don't think that emergence is fully understood yet. There are subject-specific models, but for far broader subject areas than you suggested, like coding or math or prose.
I'm sure composability is possible in principle, I'm just sceptical that it'll be a good long-term solution, for my originally stated reason. It's basically just The Bitter Lesson again, we may gain some short-lived edge by putting more domain knowledge into the algorithm, but ultimately (these days often: surprisingly quickly) it'll be outgunned by something that just leverages raw computation better.
--chat-template-kwargs '{"preserve_thinking":true,"reasoning_effort":"medium"}'
[0] https://unsloth.ai/docs/models/qwen3.8#thinking--preserve-th...There are many episodes in Roman history that people would have experienced as decline or even collapse within their lifetimes, including multiple sackings of the city of Rome itself. But in most of those cases things actually did bounce back and the Empire continued to dominate. That's kind of the point I was making. Empires would seem to be far more resilient than contemporary observers would think if one only looked at Roman history. But the problem is that the Roman Empire is a bit of an outlier in that regard. Ancient cultures of Central America would be interesting to study there (there were far more than just the Mayas, Inkas and Aztecs, most of them relatively short-lived, hence good case studies for the question of collapse). Unfortunately much less is known about them than about Romans.
It all depends on what you define as the collapse of the Roman Empire, but the earliest sensible candidate date would be about four and a half centuries after the population 'opted' for authoritarianism (which wasn't much of a choice for anyone btw).
People need to stop making senseless historical comparisons.
That's a completely normal evolution for any succesful influencer. In what world is that retirement? The only thing I can see being retired is his old online persona, the new persona seems more mature, likely tailored to please his fanbase who presumably have grown out of their teenage instincts (and have more money to spend now).
tl;dr: I don't think you and I have the same meaning associated with the word 'retired'.
My perspective is more like a FOSS philosophy for math. Even if a closed version has the same immediate effect, it's just better for everyone if everyone can look under the hood and tinker with it.
It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.