107 karma · joined January 14, 2026
I find that running better quantization, like Q8 tend to prevent this even though its a bit slower to run, it saves overall time with less churn
Using 3.6-27b is even slower again than 3.6-35b, but I find the accuracy really pays off
I'm curious about why this is
Outside of an actual test detonation, presumably this could all happen in a secure place?
Or it has been, and cruelty is the point
I can't help but get the feeling you have use-case end-goal in mind that's opaque to many of us who are gpu-ignorant.
It could be helpful if there were an example of the type of application that would be nicer to express through your abstractions.
(I think what you've shown so far is super cool btw)
That's a big assumption. Often there's no time to do things right, or no money, or lack of oversight, and so on.
Not every company is staffed by empowered and highly motivated staff
There is no such requirement.
Wouldn't the assumption be the opposite, in that AI is magnifying the decision making of the engineer and so you get more payback by having the senior drive the AI?
No they need to tariff/ban things that are non-EU