HNHacker News
TopNewBestAskShowJobs

davrosthedalek

3,105 karma · joined August 12, 2012

submissionscomments
davrosthedalek··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Exactly. And that labeled data could be collected by just recording what they feed jev and what the decision is.
davrosthedalek··on Show HN: Panda, the world's first personal AI computer
Really needs some details on what hardware is in there / what models it can run.
davrosthedalek··on Claude Opus 5.5
Well, I guess it's "fast paced".
davrosthedalek··on Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
The price is that you essentially glue CPU and GPU together, which limits total compute, from a size and thermal perspective.

This is really not a limit because of unified memory -- in principle, PCIe GPUs could read/write main memory without the CPU. But it's a limit for /fast/ unified memory, because fast means close.

So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages: a) probably faster transfer CPU<->GPU (but that's an implementation choice for the non-unified case b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.

davrosthedalek··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
regarding my brain, my mother might disagree on the laundry part.
davrosthedalek··on Nvidia agrees to acquire Hugging Face for $13B
And I would say it's actually a more secure future demand channel than the hyperscalers. Openai and co can build their own chips, so they won't buy from NVidia. But they probably won't sell them to somebody else, so everyone who wants to run their own install (and can, not the least because of HF), will buy Nvidia.
davrosthedalek··on OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)
Because it breaks the TOS.
davrosthedalek··on Our position on open-weights models
Because of regulation, not because they are intrinsically hard to get.
davrosthedalek··on Our position on open-weights models
It is actually an interesting conundrum.

Is a non-well-aligned frontier level AI a problem? I think it is likely that it is, or at least has a high likelihood to be in the future. Two scenarios for this: Misused by some bad guys. Or the terminator scenario. Both not great.

So what do we do about it?

1) We can accept it, and hope that the good guys AI can defend.

2) We can try to limit the access to it (AI proliferation?)

3) We stop the development of it

4) We can accept the risk and do nothing.

None are particular good options. Really reminds me of nuclear proliferation, on so many levels. For that, we kinda do all three:

1) Nuclear triad / iron dome / early warning systems

2) Nuclear anti-proliferation treaties.

3) Dead Physicists

Ok, so assuming all of this is true, open weights are a problem. Don't get me wrong, I love open science, open source etc. It's great to have access to capable open models. But: Even if release open weights are well aligned and have a safety layer built in, it is likely not to difficult to abliterate that part of it.

If this is really where it is going, then even closed weight model providers will see a lot more requirements for protection of the weights.

davrosthedalek··on OpenAI and Hugging Face address security incident during model evaluation
Well, I think it's fair to assume that a) They didn't upload everything before they realized it would work. b) They want to mention this as an advantage for future analyses c) Even if you assume that the attack exfiltrated everything until proved otherwise, you shouldn't just disseminate all the private information, because maybe the attack didn't.
davrosthedalek··on OpenAI and Hugging Face address security incident during model evaluation
To be fair, there are really three threats:

a) People do bad stuff because LLM told them a wrong thing. Example: AI told me I should treat my heart attack by putting a fork in the outlet. Maybe similar to seeking medical advice on reddit?

b) People use LLM to do bad stuff. Example: People use LLMs to find 0 days. Get cooking recipes for poison. Write better phishing letters. This has parallels to the gun legislation question.

c) LLMs do bad stuff on their own, beyond what the people that use it intended. The case at hand might be an example of this. Maybe similar to having an animal as a pet. We will see if it's more like a house cat, lion, or black plague.

davrosthedalek··on OpenAI and Hugging Face address security incident during model evaluation
If there is an electrical connection between these downstream boxes and the inference servers beyond the power connection, it does stretch the definition of air gap.
davrosthedalek··on OpenAI and Hugging Face address security incident during model evaluation
I am somewhat more worried about a Darkstar AI.
davrosthedalek··on Meta's AI models are powering the first wave of Genesis Mission projects
Too many spaces around the --. That's not load-bearing.
davrosthedalek··on Meta's AI models are powering the first wave of Genesis Mission projects
The grant size for phase 1 is 750k / 9 month. That is actually quite substantial.

[Despite that, I generally think more money should go to science, all around. But I have COI here.]

davrosthedalek··on SpaceX wants to launch 100k more Starlink satellites for 100x the bandwidth
Lightspeed in air is close to c_vacuum. Light speed in fiber is roughly 2/3 c_vacuum. So for transatlantic it might be faster.
davrosthedalek··on US seeks cheaper hunter-killer drones after Iran destroys $1B worth of Reapers
True, but I think the US requirements are indeed different. Ukraine must repel invading forces in their own country -> lowish range, mass produced, not necessarily precision strike.

US want to project power far away from its shores -> long range, precision strike, long loitering time.

davrosthedalek··on Delta flight hit by firework while landing at Midway Airport on Fourth of July
Das Auge ist (man) mit.
davrosthedalek··on Framework's 10G Ethernet module exposes USB-C's complexity
10 Gig ethernet is 10GBps usable rate (before packet overhead). The line rates are higher to accommodate this. For 10GBase-R, it's typically 10.3125 GBps, with a 64/66 encoding. For 10GBase-T, it's 4 lanes with PAM-16 at 800 MBaud -> 12.8 Gbps raw.
davrosthedalek··on Midjourney Medical
The question is: If you have enough full body scans of many healthy people, and the statistical tools to model it (beyond "this range is OK"), whether this would reduce these false alarms to an acceptable level.

The real crux of it remains though: Let's say it finds something that increases your death risk by x=0.1%. Could you sleep? I'm not sure. Let's say the operation has 2x=0.2% risk. What do you do? What value of x makes this a problem for you?

davrosthedalek··on Midjourney Medical
Just for illustration: Gravitational wave detection is on the femtometer scale. The proton is about that size. We can measure these things, but the machines are, let's say, "big".
davrosthedalek··on Midjourney Medical
This is somewhat speculative, but as I see it, there are two ways to retain excellent people:

a) You pay them handsomely

b) You do shit they like, they way the like.

Sometimes it overlaps, of course. But this is essentially the reason why people stay in academia in the hard sciences. Most of us could earn considerably more in industry.

I'm not sure midjourney can compete with the bigwigs on a). But doing healthcare stuff is probably more fulfilling to the researchers, and with less "we stole from all the artists" vibes.

Of course, if this all works out, they might me able to do a) easily :)

davrosthedalek··on Did Claude increase bugs in rsync?
Right, that's now /usr/bin/true !
davrosthedalek··on Did Claude increase bugs in rsync?
Why it is probably and regrettably true that few people in camps will change their mind, data analysis can help people who haven't been captured yet to either stay away from the camps or at least fall into the "more correct" one.

Stay out of camps, people!

davrosthedalek··on Did Claude increase bugs in rsync?
First rsync and now less? What comes next, cat?
davrosthedalek··on Did Claude increase bugs in rsync?
But: "After posting this on Hacker News and recieving [sic] almost no substantive input, discussion, or response on the actual content of the article, I decided to rewrite all of the prose in my own voice. If anyone complains about my verbosity or sentence structure — as they usually do, which is the reason I originally let the AI write the prose, among other reasons obsoleted by templating — they can go fuck themselves."

So rewritten in his own voice. Maybe the m-dashes are from GLM, maybe from the author.

davrosthedalek··on Did Claude increase bugs in rsync?
Haven't looked at the code, but the allocated memory could be larger than necessary to make "off-by-one" or "off-by-a-few" errors less deadly. Then zeroing it out makes it even less so. Defense in depth.

Or it's an allocation for an arena? The zeroing might help trigger 0 derefs earlier if the overrun happens for the object that are then allocated in the arena (and not by allocating more objects than the arena can provide)

davrosthedalek··on Did Claude increase bugs in rsync?
Many statistics were presented. In the view of the author (and I think he is correct), none of them show evidence for an increased bug rate from Claude. That is absence of evidence (...for the increased bug rate).

The two examples you bring are not claims of absence of evidence, but claims of evidence of absence. The author takes the result as evidence that there is no effect. As I wrote, the author shouldn't do that, because indeed you cannot distinguish between "no effect exists" and "no effect observed". But again, these are (wrong) claims for evidence of absence.

The author can absolutely claim: I did these statistical tests, and none showed evidence that there is an effect. Absence of evidence. It's not a claim that there will never be evidence. Just that there is none from these tests.

Edit: To convert the absence of evidence into evidence for absence, indeed you need to understand the statistical power of your test, and how it is affected by alternate hypotheses. And for that, without having done the math, having only two data points seems very thin.

davrosthedalek··on Did Claude increase bugs in rsync?
Yes, I do for example.

And the author discussed the use of AI pretty exhaustively in point 0 of the post.

davrosthedalek··on Did Claude increase bugs in rsync?
No. It's a description of the result of the maybe underpowered study. the underpowered study did not find evidence. Evidence is absent. Because it is underpowered, it's not evidence that the effect is absent.

The claim is not "two experimental conditions did not differ". The claim is "The data do not show evidence that the experimental conditions did differ".

Page 1 of 34Next →