HNHacker News
TopNewBestAskShowJobs

gojomo

30,203 karma · joined February 24, 2007

Gordon Mohr, maker of software.

Twitter: http://twitter.com/gojomo

Project: http://thunkpedia.org

Idea blog: http://memesteading.com

Older blog: http://gojomo.blogspot.com

Homepage: http://xavvy.com (username @ here for email contact)

My HN peeve is formulaic downbeat comments, like: "How is this news, I already knew this!" "…Betteridge's Law…" "I stopped reading at…"

submissionscomments
gojomo··on A computer upgrade shut down BART
Pre-opening BART tubes were definitely used for George Lucas's first feature film, THX-1138: https://www.sfgate.com/streaming/article/bart-transbay-tube-...
gojomo··on Study mode
Grandparent testimony of success, & parent testimony of frustration, are both just wispy random gossip when they don't specify which LLMs delivered the reported experiences.

The quality varies wildly across models & versions.

With humans, the statement "my tutor was great" and "my tutor was awful" reflect very little on "tutoring" in general, and are barely even responses to each other withou more specificity about the quality of tutor involved.

Same with AI models.

gojomo··on GLM-4.5: Reasoning, Coding, and Agentic Abililties
Given that other work shows that models often converge on similar internal representations, I'd not be surprised if there were close analogues of 'subliminal learning' that don't require shared-ancestor-base-model, just enough overlap in training material.

Further, "enough" training from another model's outputs – de facto 'distillation' – is likely to have similar effects as starting from a common base model, just "from thge other direction".

(Finally: some of the more nationalistic-paranoid observers seem to think Chinese labs have relied on exfiltrated weights from US entities. I don't personally think that'd be a likely or necessary contributor to Z.ai & others' successes, the mere appearance of this occasional "I am Claude" answer is sure to fuel further armchair belief in those theories.)

gojomo··on Ask HN: Is it time to fork HN into AI/LLM and "Everything else/other?"
I'd expect a noticeable delay with current local LLMs - especially visiting a site for the 1st time. But then they could potentially memoize their heuristics for certain designs, including recognzing when some "deeper thought" newly required by server-side redesigns.

But of course local GPU processing power, & optimizations for LLM-like tools, all adancing rapidly. And these local agents could potentially even outsource tough decisions to heavierweight remote services. Essentially, they'd maintain/reauthor your "custom extension", themselves using other models, as necessary.

And forward-thinking sites might try to make that process easier, with special APIs/docs/recipe-interchanges for all users' agents to share their progress on popular needs.

gojomo··on Ask HN: Is it time to fork HN into AI/LLM and "Everything else/other?"
~simonw's demo of a quickie customized HN front-end is great.

But ultimately, your browser should have a local, open-source, user-loyal LLM that's able to accept human-language descriptions of how you'd like your view of some or all sites to change, and just like old Greasemonkey scripts or special-purpose extensions, it'd just do it, in the DOM.

Then instead of needing to raise this issue via an "Ask HN", you'd just tell your browser: "when I visit HN, hide all the AI/LLM posts".

gojomo··on Ask HN: Is it time to fork HN into AI/LLM and "Everything else/other?"
https://en.wikipedia.org/wiki/A_Void
gojomo··on Measuring the impact of AI on experienced open-source developer productivity
Thanks, that's great!

But: if all developers did 136 AI-assisted issues, why only analyze excluding the 1st 8, rather than, say, the first 68 (half)?

gojomo··on Measuring the Impact of AI on Experienced Open-Source Developer Productivity
Did each developer do a large enough mix of AI/non-AI tasks, in varying orders, that you have any hints in your data whether the "AI penalty" grew or shrunk over time?
gojomo··on Public Signal Backups Testing
It doesn't. Signal opts its data out of even encrypted local iPhone backups.
gojomo··on Public Signal Backups Testing
And note: Signal has been misleading users about options for iOS backups for a decade! In September 2015, users were told it was "on the roadmap": https://github.com/signalapp/Signal-iOS/issues/905#issuecomm...

No option has ever existed on iOS, despite the recent announcement assuring users "Local backups still exist". They've never existed for iOS users!

gojomo··on Public Signal Backups Testing
Do such backups retain Signal message history - unlike Apple's native local backups to MacOS?
gojomo··on Public Signal Backups Testing
News to me! Where is this option described, ideally by Signal itself?

If you are alleging that Apple's own local Finder/Itunes backup of an iPhone includes Signal messages, that's not true, against reasonable user expectations, by Signal's own design choices.

Anyone who's counting on such local backups to save their histories is in for the same rude surprise I and many others have hit unaware:

https://www.reddit.com/r/signal/comments/1hgukpg/backup_and_...

gojomo··on Public Signal Backups Testing
The post claims with regard to the cost of a cloud backup that "Local backups still exist" - but that's a lie, there's no local backup option on iOS.
gojomo··on Unsupervised Elicitation of Language Models
> far from novel

Techniques can be arbitrarily old & common in industry, but still be a novel academic paper, first to document & evaluate key aspects in that separate (& often lagging) canon.

gojomo··on Low-background Steel: content without AI contamination
No - and you can compare the style & written tics for continuity with my 18y of posts here.

I used 'delving' in an HN comment more than a decade before LLMs became a thing!

https://news.ycombinator.com/item?id=1278663

gojomo··on Low-background Steel: content without AI contamination
Yes, but: for humans, even without an expert-over-the-shoulder providing fresh feedback, drilling/practice works – with the right caveats.

And, counter to much intuition & forum folklore, it works for AI models, too – with analogous caveats.

gojomo··on Low-background Steel: content without AI contamination
Of course, training on synthetic data can't do everything! My main point is: it's been doing a bunch of surprisingly-beneficial things, contra the obsolete beliefs about model-output-worthlessness (or deleteriousness!) for further training to which I was initially responding.

But also: with regard to claims about what models "can't experience", such claims are pretty contingent on transient conditions, and expiring fast.

To your examples: despite their variety, most if not all could soon have useful answers answers collected by largely-automated processes.

People will comment publicly about the "vibe" & "people-watching" – or it'll be estimable from their shared photos. (Or even: personally-archived life-stream data.) People will describe the banana bread taste to each other, in ways that may also be shared with AI models.

Official info on policies, processing time, and staffing may already be public records with required availability; recent revisions & practical variances will often be a matter of public discussion.

To the extent all your examples are questions expressed in natural-language text, they will quite often be asked, and answered, in places where third parties – humans and AI models – can learn the answers.

Wearable devices, too, will keep shrinking the gap between things any human is able to see/hear (and maybe even feel/taste/smell) and that which will be logged digitally for wider consultation.

gojomo··on Low-background Steel: content without AI contamination
OK, sure, there are gradations.

The new encoding can contain a FLOAT32 side channel on every character, to represent its proportional "AI-ness" – kinda like the 'alpha' transparency channel on pixels.

gojomo··on Low-background Steel: content without AI contamination
Perhaps. But these models can already clearly write about the world, in useful ways, without such 'qualia' or 'biological underpinnings'.
gojomo··on Low-background Steel: content without AI contamination
No, new more-capable and/or efficient models have been forged using bulk outputs of other models as training data.

These inproved models do some valuable things better & cheaper than the models, or ensembles of models, that generated their training data. So you could not "just ask" the upstream models. The benefits emerge from further bulk training on well-selected synthetic data from the upstream models.

Yes, it's counterintuitive! That's why it's worth paying attention to, & describing accurately, rather than remaining stuck repeating obsolete folk misunderstandings.

gojomo··on Low-background Steel: content without AI contamination
The training sets can already include direct data series about the world, where the "work of human beings" is just setting up the the collection devices. So models can absolutely "experience the world".

But I'm not suggesting they'll advance much, in the near term, without any human-authored training data.

I'm just pointing out the cold hard fact that lots of recent breakthroughs came via training on synthetic data - text prompted by, generated by, & selected by other AI models.

That practice has now generated a bunch of notable wins in model capabilities – contra the upthread post's sweeping & confident wrongness alleging "Ai generated content is inherently a regression to the mean and harms both training and human utility".

gojomo··on Low-background Steel: content without AI contamination
Less than you might think! Some of the frontier-advancing training-on-model-outputs ('synthetic data') work just uses other models & automated-checkers to select suitable prompts and desirable subsets of generations.

I find it (very) vaguely like how a person can improve at a sport or an instrument without an expert guiding them through every step up, just by drilling certain behaviors in an adequately-proper way. Training on synthetic data somehow seems to extract a similar iterative improvement in certain directions, without requiring any more natural data. It's somehow succeeding in using more compute to refine yet more value from the original non-synthetic-training-data's entropy.

gojomo··on Low-background Steel: content without AI contamination
This was an intuitively-appealing belief, even with some qualified experimental support, as of a few years ago.

However, since then, a bunch of capability breakthroughs from (well-curated) AI generations has definitively disproven it.

gojomo··on Low-background Steel: content without AI contamination
Look, we just need to add some new 'planes' to Unicode - that mirror all communicatively-useful characters, but with extra state bits for...

guaranteed human output - anyone who emits text in these ranges that was AI generated, rather than artisanally human-composed, goes straight to jail.

for human eyes only - anyone who lets any AI train on, or even consider, any text in these ranges goes straight to jail. Fnord, "that doesn't look like anything to me".

admittedly AI generated - all AI output must use these ranges as disclosure, or – you guessed it - those pretending otherwise go straight to jail.

Of course, all the ranges generate visually-indistinguishable homoglyphs, so it's a strictly-software-mediated quasi-covert channel for fair disclosure.

When you cut & paste text from various sources, the provenance comes with it via the subtle character encoding differences.

I am only (1 - epsilon) joking.

gojomo··on Meta: Shut down your invasive AI Discover feed
Lacks context & examples to know what they're concerned about.

Has a righteous, bossy tone that doesn't seem earned by case particulars or its (anonymous) author.

"Mozilla: Improve your messaging. Now."

gojomo··on Company Reminder for Everyone to Talk Nicely About the Giant Plagiarism Machine
There is no de jure legal requirement that the RIAA, Disney, Nintendo, or the government be "pleased to hear" about new technology.

And, while copyright prohibits some sorts of reproduction of copyrighted materials, it doesn't give rightsholders veto power over all downstream uses of legal copies.

gojomo··on ChatGPT could never get a PhD in geography
The giveaway that Marcus is generating slop for consumption by unsophisticated AI-haters is that he doesn't bother to mention what models/options involved.
gojomo··on Show HN: Undetectag, track stolen items with AirTag
After I saw third-party "10 year battery enclosure" offerings for AirTags, was wondering when other workaround customizations like this might appear.

Other impactful variants might be:

* senses whether another 'sibling' AirTag is present, if so, stays off. If not, waits X hours & then turns on.

* has its own motion sensor; only after X minutes of being stationary, it waits Y hours to turn on briefly

* has its own clock & (original-user-known) randomization seed; turns on at pseudorandom intervals the original user can predict

* low-power/low-bandwidth receivers so cheap & tiny now: could wait for national or even global unit-specific 'wake' request - perhaps even with parameters for duration/intervals – before powering-on AirTag portion

gojomo··on Launch HN: Tinfoil (YC X25): Verifiable Privacy for Cloud AI
Is there a frozen client that someone could audit for assurance, then repeatedly use with your TEE-hosted backend?

If instead users must use your web-served client code each time, you could subtly alter that over time or per-user, in ways unlikely to be detected by casual users – who'd then again be required to trust you (Tinfoil), rather than the goal on only having to trust the design & chip-manufacturer.

gojomo··on Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs
Prior discussion when the paper was 1st reported in February: https://news.ycombinator.com/item?id=43176553
← PreviousPage 3 of 34Next →