32,237 karma · joined January 3, 2008
Surely between the squeeze on RAM and storage pricing and the exorbitant costs of training or mass-inference scale GPUs (if you’re going into AI), the nominal cost of the CPU licensing is really not going to make or brake anything or unlock some new business viability?
Even RISC-V aside, we are essentially in a position that would have been impossible to even dream of ten or twenty years ago when Intel was the only player in the game.
Wouldn’t buying off-lease hardware (not even in bulk) give you better performance per dollar, better compatibility, and more options?
Not to say any of this will always be the case; even if RAM pricing doesn’t come down, storage will, and RISC-V processors will (maybe) eventually actually be competitive when it comes to SOTA performance, but today in 2026?
Interesting times, to say the least!
The comparison should be against renting in the cloud for the duration of your task for training and research or using pay-per-api-call providers for general inference instead of buying your own hardware (and paying the electricity and cooling bills on top), because let’s face it, the models you want to use are probably the same ones available on inference providers (but, yes, some are more trustworthy than others).
Speaking as someone that does ML/AI research, you are essentially paying a huge premium for being able to just run your Python script at any time without setting up a deployment script and harness to run the job remotely, while your hardware sits essentially idle the rest of the time.
The only way to make the math work is if you rent your hardware in the background for inference while you’re not using it in anger, but despite all the startups and promises that has never become as streamlined as mining bitcoins or shitcoins used to be and they don’t pay out as much as they say they would. Renting your hardware for training is another option but doing that is a lot more involved, options are fewer and farther in between, you won’t get as much utilization out of it, and doesn’t let you feasibly abort running tasks at a moment’s notice.
Is the JpegXL lossless options the "transparent JPEG recompression" or the actual lossless profile? I'm presuming the latter because it more than doubled the image size.
The problem becomes when you add in the adjustable reasoning efforts and you end up with {model, reasoning_effort} combinations that end up completely obviating particular model classes altogether for at least some percentage of queries; e.g. with GPT 5.6 the price/performance Pareto frontier is dominated by permutations of either Luna and Sol, with Terra nowhere to be seen (but then if you need "large model smells" that aren't captured by your benchmark you can't even rely on this, as a model like Luna simply isn't capable of encoding sufficient world knowledge in its weights to perform certain tasks at any reasoning level but you might be able to get away with Terra on low reasoning, but no one seems to be covering this for some reason).
Model Input Cached input Cache writes Output
gpt-5.6-sol $4.00 $0.40 $5.00 $20.00
gpt-5.6-terra
$2.00 $0.20 $2.50 $12.00
gpt-5.6-luna
$0.20 $0.02 $0.25 $1.20
So Sol is still 20x Luna, but much more appealing when compared to offerings from Anthropic and others.Obvious next step is to explore if you can replace watermarker.dll with a (signed) no-op shim or MITM the API call to at least use your own (nil?) GUID that isn't linked to your device/account.
In case it's not obvious, my bigger concern isn't "this image can be identified to have been generated with/by AI" so much as it is "digital yellow printer dots have been forced upon us, except they can identify and retrieve the exact user/device/time/place/document/etc", completely destroying any and all illusions of privacy left.
3.0 flash (not lite) handled it like a champ though, fwiw.
I appreciate the article nevertheless, of course, but I do feel that it would probably be more meaningful to someone that has at least a basic understanding of Japanese.
In fact, (semantics aside, from a technical perspective) the preference should always be for modifiers rather than standalone characters because the chances of being supported by the viewer’s font are much greater: it doesn’t need a separate glyph explicitly drawn and add to the font file for the code point. Difficulties in entering it or typing it out should be mitigated with client-side affordances in the UI, shortcuts, etc.
For example, I mocked a dumbed down version of what would be a reasonable intermediate tool call prompt:
> It's currently 58 degrees. User asks for house to be 8.5 degrees warmer. What temperature to set thermostat to?
The reply?
Reasoning: “User asks for temperature to set thermostat to 8.5 -> set_thermostat with temperature=8.5.”
Sounds like something Siri would do!