HNHacker News
TopNewBestAskShowJobs

myworkaccount2

13 karma · joined April 14, 2025

submissionscomments
myworkaccount2··on DeepSeek V4 Pro 0813
IMO the HLE scores without tools seem to align better with real world performance of the models.

To me it feels like the difference between "RL performance" and the pretraining / base "knowledge".

Yes you can RL terminal bench to the moon but does the model hold up on out of distribution tasks?

Kind of like trying to navigate a dark room with a laser light, vs a flashlight. Laser is going to go a lot farther much more efficiently but only if you are already pointing it at the right place.

myworkaccount2··on Stealing Reasoning Traces from Proprietary LLM APIs
There seems to be an obvious choice to make here, should you give the users to decrypt and use the COT that they did not generate themselves?

This is only required if you want users to be able to share things with everyone and you are going for the simplest implementation.

If not you could try to keep a record of keys associated with a user, then when a new request comes in look through to see if the user has a valid key to decrypt the COT.

For explicit shares, just add the key used in that one conversation to the users valid keys. For global shares use the global keys. But that's adding more complexity to the system.

myworkaccount2··on Stealing Reasoning Traces from Proprietary LLM APIs
Is this how the eastern labs "distill" SOTA models?

If you can play it right, you don't even need to send suspicious prompts to the frontier models. Just use them for regular tasks, extract the encrypted COT blocks and replay it to a cheaper model to get the plain text COT.

But the real question is: Is it okay to steal from a thief's hoard?

myworkaccount2··on Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines
I wonder if this kind of analysis will give us a way to check if the frontier labs are waiting for the right moment to release their models. To me it feels obvious that these companies are not releasing models as soon as they are done doing their post training / testing with any new model.

But there is no real way to know how much of this "waiting" any lab is doing, if we can get better estimates this way maybe we can gauge how far the open weights models really are.

myworkaccount2··on CrankGPT
I can't tell if this is a joke, or they are serious.
myworkaccount2··on Claude Opus 4.8
Anyone else experiencing tool call failures? Switch back to 4.7, same prompt, same everything it works with no problems.
myworkaccount2··on Claude Code wiped our production database with a Terraform command
So even if you delete everything and make sure to keep no backups, amazon can still recover the db. What am i missing here?
myworkaccount2··on Show HN: DoNotNotify – Log and intelligently block notifications on Android
when will it be available through fdroid?