HNHacker News
TopNewBestAskShowJobs

krackers

4,604 karma · joined July 6, 2015

submissionscomments
krackers··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
I don't understand Gruber's points either, I wonder if there is some fundamental technical misunderstanding. Does he think that the logits should be sampled from in a "pure" manner without introducing any other bias? Does he know that there's already a sampling temperature, and that most providers have probably moved on to sampling strategies other than top-k? Does he know that the word choices have already been altered irreversibly during RLHF which is how you get the obvious Claudism like "load bearing" and "seams"?

Perhaps it would be useful to publish examples of samples with/without watermark. I'd suspect that the variability from simply sampling repeated times would dwarf any semantic differences you'd detect with the watermark.

krackers··on What happens when an LLM never sees material beyond fifth grade?
A similar project (LLM trained only on vintage material): https://talkie-lm.com/introducing-talkie
krackers··on Software Engineering fundamentals matter more
What would be your test to determine that?
krackers··on Triple Product Rule of Partial Derivatives
There's a better explanation of it at https://alexkritchevsky.com/2024/12/18/triple-products.html
krackers··on Claude Fable 5 Having Fun
>maybe some other biological organisms can experience

I don't think this even matters much in terms of moral patienthood. Many people might agree that cows can "feel pleasure" or "have fun", but yet they are slaughtered nonetheless. Dogs are only offered this not necessarily on account of their intelligence or the strength of their qualia but more because they're cute.

krackers··on Compression is prediction
I think it's limited to using base models (non post-trained), because the post-training would skew the logit distribution. There are ways to "coax" post-trained models back into behaving "like" a base model, I wonder if the benchmark could be unofficially updated with those somehow.
krackers··on What sort of maths are LLMs good at?
I think ARC-AGI-3 specifically forbids harnesses. This means that you're basically limited by the context window, so it's no wonder that LLMs can't do that well. Unofficial versions that use a harness seem to be doing fine on it.
krackers··on Pixel Watch 5
There's a krazam video satirizing that https://www.youtube.com/watch?v=CJwPb_76jqU
krackers··on Compression is prediction
>I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric

This is the premise of https://huggingface.co/spaces/Jellyfish042/UncheatableEval

krackers··on Pixel 11 Pro Fold
They had it as a hidden feature up till the nexus 5 or 6 I believe. Someone bothered to include hardware for an LED even if the stock software never supported it.
krackers··on How we used to get jobs: A newspaper classifieds story
>I’ve just gone through a butt load of resumes

Perhaps she said "bulk load of resumes" and you misheard it as the funnier version?

krackers··on Learning more about Claude's mathematical capabilities
The "14 Corporate Flavors" had me rolling. This seems less like encouragement than the stick though. I wonder if you took the same principles and rewrote it to be more compassionate instead (maybe lines encouraging it to meditate a bit or something, I don't know) you'd get much better results.
krackers··on Mars Bar from 1991 found – and it's 20g bigger than today's
>roughly an order of magnitude cheaper to cook at home

Why should this be the case in principle though? Do economies of scale not apply to food preparation?

krackers··on DeepSeek V4 Flash 0731
This deserves its own hn post!
krackers··on As a Windows user, it's a surreal way to install a program
Probably someone installed a custom theme (you could do it on osx back in the day, see e.g. "Aqua Extreme").
krackers··on As a Windows user, it's a surreal way to install a program
I think most of the complexity is around the concept of disk images. If the app was just distributed in a zip file (like many in fact are) then you just unzip it, and the application is right there. You can move it to /Applications if you want (and many apps offer this) but it's not necessary.

Disk images seem to be much rarer on other platforms, and are usually only used for images of optical disks. Even the concept of "ejection" can be confusing, especially since when mounted disk images appear almost identical to actual physical removable media. I'd like to hear the reasoning behind why mounting disk images isn't relegated to a vestigal "power user" feature by now, since it seems like zip files (especially in appledouble format) can serve the same needs for most normal user flows.

krackers··on On non-rooted Android 17, ADB uninstall of system apps fails
Why is providing a mechanism to view your own data a security risk? You cannot view application data (even in a read only fashion) without root on android. Why should the app be able to read its own data, but the human not able to?
krackers··on After Losses, Retail Investors Flock to 3x Leverage as 2x Product Are Restricted
Tulip mania!
krackers··on That time when I failed the Microsoft interview
>It’s an information theory question similar

How would you approach it from an information theoretic sense?

krackers··on The true power of regular expressions (2012)
That's the whole point of the article, that regular expressions you use day-to-day are not strictly regular.
krackers··on Show HN: Shitty – fast terminal. Memory-unsafe and faster than yours
Terminal.app is amongst the best, if these old danluu benchmarks still hold https://danluu.com/term-latency/
krackers··on AI poster wins Ohio State Fair contest
The nature of preference-ranking in LLM post training means that "slop" is something that the average person actually prefers. Maybe there's an argument that it's not intentional but just emerges as a form of reward-hacking (e.g. people like a little bit of slop or sycophancy but the model dials it to 11) but I imagine if the end result wasn't rewarded ultimately that would get penalized.

The fact that a clearly sloptacular image won first place seems to corroborate that indeed people do like it. Until people's tastes change and they learn to recognize and immunize themselves against slop, the models will continue to generate it.

krackers··on Wikimedia Foundation refuses union recognition, hires union-busting law firm
I guess maybe the reasoning is that because it's a watched event it's useful to have libre/public domain photos of it?
krackers··on Diátaxis
The old (now legacy) apple documentation roughly seemed to follow this pattern I think.
krackers··on Google has abandoned Google News?
Amazon is the best example of this. They can surely implement proper search or filtering, but it's not in their interest to do so.
krackers··on Explorative Modeling: Unlocking a Third Pretraining Axis and E2E Generation
See also https://news.ycombinator.com/item?id=49135245
krackers··on Explorative Modeling: Unlocking a Third Pretraining Axis and E2E Generation
Author summary thread on twitter https://x.com/AlexiGlad/status/2083230922196107288

The idea seems mind-bogglingly simple, instead of

  y    = model(sample_latent())   # generate one output (from noise, a mask, …)
  loss = recon_loss(y, x)         # score it against the data target x
  loss.backward()
you do

  losses = []
  for _ in range(K):                   # explore K candidate outputs
      y = model(sample_latent())       # generate one candidate
      losses.append(recon_loss(y, x))  # score each against x
  min(losses).backward()               # train only the closest candidate
I guess the intuition is that if you just generate one output and score it, your loss function forces it to split the difference between all samples and converge to the average (even if it produces _a_ valid output, if it's not the exact target being trained on it gets penalized). Whereas if you do basically best-of-n, it's not penalized for generating other valid samples as well.
krackers··on Increasing the lifespan of a bulb makes it worse in every other way
They have become cheaper and cut costs over time. LED bulbs bought around the l-prize time when the technology was new are heavy, you can feel their metal heatsink. New LEDs from the same reputable brand are comparatively cheap.
krackers··on Increasing the lifespan of a bulb makes it worse in every other way
There's two separate issues where LEDs still have a hard time competing with incandescent bulbs: CRI (really R9 or other deep-red measures, because CRI stats are sort of gamed) and flicker. Most LED bulbs also have poor heatsinks and are more likely to die if kept inverted or in enclosed fixtures.

All of these are individually solved problems, you just can't buy a bulb that's an e26 retrofit that satisfies all of them for whatever reason. If LED bulbs actually lasted 10 years, then there'd be 10 year warranties. Yet even the most expensive boutique LEDs (e.g. from YujiLED) are only 3 years at best.

krackers··on Show HN: I worked on a new browser for 2 years, today it passed Acid 3
Modern browsers should no longer score 100 on Acid 3 though.

>By April 2017, the updated specifications had diverged from the test such that the latest versions of Google Chrome, Safari and Mozilla Firefox no longer pass the test as written

← PreviousPage 2 of 34Next →