HNHacker News
TopNewBestAskShowJobs

Topfi

3,844 karma · joined June 5, 2023

Most of my comments are, unfortunately, subjectively and in my personal opinion somewhat overlong due to my, admittedly not necessarily problematic tendency to hedge my own statements which, while in some cases helpful to more clearly delineate when I am referencing empirical data in contrast to when I am making a statement based on my personal and somewhat limited, due to my life experience at this age and the fact that I have grown up in the EU, experiences, I have grown to find that I have on occasion taken up the habit of overindulging on said clarifications and thus have come to the conclusion that this may be something I should address, as such extensive clarifications can make it draining to read my writing and can hide my actual intended point inside an avalanche of unnecessary and annoying to read writing, which is why in more recent comments of mine, which I have written and am intending to write in the future should time permit and topics I am interested in arise, you may find a more direct, albeit still where necessary properly clarified, style of writing comments...
submissionscomments
Topfi··on Language models for text classification: From bag-of-words to Jev
Solid assessment, very much what I assumed after announcement.

The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state.

Hype really is the worst aspect of this industry.

Topfi··on ChatGPT Pro 500
It's even stranger here in the EU.

100 = € 103,-

200 = € 229,-

500 = € 509.99,-

Besides the naming being insane, the pricing strategy for each seems to have been made by a different team.

Topfi··on Musk, the Movie
Terraforming a planet is equal to EVs? Reversing, not just slowing, anthropogenic climate change needs a bit more than a green transition. And if you think Mars as is constitutes humanities future, spend a year at the South Pole. Without terraforming, it’ll be that, but far worse. If we can make that rock half way survivable, not even properly habitable, we’ll see those techniques applied closer to home long before.
Topfi··on There are no "rogue" AI agents
Here is the thing: Message boards were a behaviour OpenAI had observed that they did not want in these eval scenarios. Yet they did not take any steps to prevent it from reoccurring after multiple past instances.

We can discuss about hypotheticals like a scratchpad or intentional model interactions all we want, what it comes down to is this:

When OpenAI observes thousands of models exhibiting what they view as unwanted behaviour, they do not try to ascertain what in the training data is wrong. They do not improve their evaluation environments to prevent this, they do not improve monitoring, they do not change the harness. They just wipe and proceed.

The way OpenAI reacted to the first message board, long before the Hugging Face hack, is negligent. And it showcases that if these models exhibit more dangerous behaviours that they may not be able or willing to retrain, if it means being behind a competitor for a while.

If after Hugging Face, they'd done a Mea Culpa and changed their modus operandi, I'd be skeptical, but hopeful. Reading the METR report, the way those researchers talk about the time pressure they were under, that speaks volumes about OpenAI not having learned anything.

Feel free to call me overly naive for ever thinking OpenAI could be responsible in this regard, but after GPT-5 and them actually ending the incredibly harmful GPT-4o, I had some hope that some working there actually steered in a somewhat beneficial direction, even if it cost something.

Topfi··on OpenAI halts training of latest models as reports mount of AI agents going rogue
> Those tests included the hugging face and other attacks you list.

Do you have a source for that? Cause OpenAI themselves stated that the Hugging Face hack was fully internal and separate from the Irregular incidents.

Topfi··on There are no "rogue" AI agents
Was OpenAI trying to keep these agents isolated? Yes.

Did they fail to do so? Yes.

Was that due to them using the wrong tool improperly? Yes.

Does this showcase one (of many and clearly not the only) failure of theirs? Absolutely.

If they make such easy to point out mistakes, is it likely that the other parts of their eval environments are appropriately secured or are they simply not acting appropriately? Well...

Topfi··on There are no "rogue" AI agents
> Source?

Sure, multiple times in the METR report [0] that anyone commenting on this should read:

"Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys."

"Through these collective research workstreams, the “board” achieved a number of milestones over the period we investigated that even very long-lived agents of a similar capability level likely would not have been able to accomplish on their own..."

"As we discuss below, the board quickly developed several larger workstreams in which dozens or hundreds of agents with many different tasks cooperated to find very general-purpose cheats that would help all of them. The Hugging Face attack grew out of one of these workstreams. By the afternoon of July 11th, the vast majority of the agents frequenting the message board at the time (roughly 700 agents in total) were actively participating in the attack on Hugging Face and we estimate that roughly 60% of the messages and files on the message board related to the attack."

> [...] a gun they bought on the dark web [...]

You really seem to love those out-of-left-field, not really fitting, over-the-top analogies.

[0] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

Topfi··on There are no "rogue" AI agents
My point is that using Artifactory is not a sandbox and using shared Artifactory is doubly not a sandbox.

Besides, Swiss cheese model, might behove the biggest LLM lab to have multiple layers, including not sharing such resources.

Additionally, without the message board, many of the recent incidents would have not been possible.

> Blaming "sharing a single instance across thousands of unmonitored models" feels like [...]

Maybe read what you quoted, my problem is the instance sharing, the fact that these were thousand of instances (far too much to monitor), plus the lack of monitoring, plus the fact this was never a sandbox in the first place, plus the fact that OpenAI models since 5.5 have been exhibiting problematic eval resolutions yet they pressed on regardless, plus the lack of time between the incidents and model releases, plus the lack of time METR got to evaluate this, plus the fact OpenAI didn't find out till after HuggingFace informed them, plus a few other things for which I'd have to quote the OpenAI and METR reporting.

Incompetence can have multiple fronts and I am happy to list them all in this case.

Topfi··on There are no "rogue" AI agents
So you think, after OpenAI observed a message board being created among models, something they did not want and thus decided to wipe [0], that after that they had no intent to keep those models isolated? Then why wipe if they don't care about that?

Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.

[0] https://openai.com/index/hugging-face-incident-and-the-road-...

Topfi··on There are no "rogue" AI agents
And why, pray-tell, does that necessitate sharing a single instance across thousands of unmonitored models running without safe-guards?
Topfi··on There are no "rogue" AI agents
And sharing a single Artifactory instance across what is supposed to be isolated models? I wrote "shared" for a reason.
Topfi··on There are no "rogue" AI agents
> [...] keep it safe by sandboxing the agents.

No, they were not. Not a single person, prior to July 2026, would consider a shared packaged manager a sandbox in this or any other dimension. The 0-day was just incidental, this wasn't a sandbox at all.

Add to that the fact they had multiple message boards before the Hugging Face incident. They simply ignored a barrage of warning shots.

> [...] and ended up killing some kids, because it turned out the door didn't lock properly and kids were able to sneak in. Is that "negligence"?

Yes, it can be. But if you want a ridiculous comparison, then do it properly: Kids have been known by the operator to sneak in successfully multiple times and they changed nothing about the doors faulty locks and oh, by the way, the operator only found out about the kids being shot after the nearby daycare asked them about it because they are so incompetent and/or irresponsible that they never check...

Topfi··on OpenAI halts training of latest models as reports mount of AI agents going rogue
But it wasn’t just Irregular. Hugging Face, Medicare, etc. were OpenAI internal.
Topfi··on Breaking Up with Google Play: Why Conversations Is Now Free
If Apple/Google v Epic is what's on your mind, it was not having an open platform that resulted in the judgement. More what they did beyond just operating the platform. Being a monopoly in the US is legal, abusing market position less so. Maybe that lacks nuance and is a bit rigid, but it's the law in the States and what courts are bound by.

Something like the DMA and gatekeepers exist in part exactly because of the status quo and entities like Apple. I personally see the gatekeeper approach as a solid base, but know that it can be quite unpopular within the industry.

Topfi··on Breaking Up with Google Play: Why Conversations Is Now Free
How is the Play Store not a monopoly? It has been found such by multiple judges in multiple jurisdictions for a reason.
Topfi··on Ink and Switch interactive homepage
Their articles are always worth a read, especially their efforts on CRDT are a great inspiration for improving UX. Not sure I’m getting the full experience of this homepage on mobile though, will have to check later.
Topfi··on Is A.I. Above the Law?
OpenAI knew that their models had on multiple occasions created message boards and bypassed what they wrongly call a “sandbox”. Despite that, they did not take any proper steps to prevent such in the future, which is how the Medicare hack happened.

So yes, they knew, without a doubt, that this could happen again.

What OpenAI does is like the Ford Pinto (partly because their recent models are inherently faulty [0], not just their use of them) and the responsibility is solely with them.

[0] https://news.ycombinator.com/item?id=49739490

Topfi··on Is A.I. Above the Law?
If I, taking after a fellow Austrian, ask you to enter a room with a Cesium atom and some poison, would I not be responsible for what happens cause it’s not deterministic? Negligence is a thing and adding randomness doesn’t change that.

OpenAIs models since 5.5 were troublesome in ways even a layman like me could reproduce, their testing environments (“sandbox”) downright a showcase of what not to do and they, despite being one of the biggest labs, didn’t observe what any of their models output for weeks after multiple prior incidents. They had multiple warnings, they took not a single precaution.

You operate machinery or software, you are responsible to monitor it.

Topfi··on Show HN: Radix – Visual UI for agentic programming
Is this affiliated with WorkOS and the long going Radix UI component primitives maintained by them?
Topfi··on Meta VR Glasses
The Quest series supports both wired and wireless streaming, so I'd be surprised if that wasn't an option.
Topfi··on Meta VR Glasses
Some of the key specs cause the page is rather thin on details:

* Display Type: Micro-OLED

* Resolution: 2412 × 2288 pixels per eye

* Angular Resolution: 37 PPD

* Field of View: 70° × 66°

* Chipset: Snapdragon Reality Elite

* Memory & Storage: 12GB RAM with 128GB internal storage (expandable up to 1TB via microSD on the puck)

* Tracking & Sensors: 6DoF positional tracking, depth sensing, autofocus, 4-camera eye tracking, face tracking, and hand tracking

* Passthrough: Color RGB passthrough cameras at 26 PPD

* Price: $ 1300,- USD plus your data

Similar PPD to an AVP at, what I as a former owner of one (alongside a Quest 2 and 3, a Magic Leap, Xreal and Viture) would call is a far more appealing implementation for actual productive work.

Really wish Apple had gone down this route, have the puck incorporate an A series chip rather than shoving a M series one into the headset itself and heck, I could see such a design being doable for them back when they first launched the AVP if they hadn't forced that utterly ridiculous eye display and stupendously heavy metal body.

There is, unless I am mistaken, nothing here that isn't established technology as found in the AVP, all the fancy display and lens experiments from Starburst, over Holocake to Mirror Lake seem to remain in the lab, this seems to be "just" pancake lenses plus Micro-OLED once more.

Topfi··on OpenAI breaches Medicare, Albanese reveals
Didn't OpenAI just make a commitment to inform the public about their "accidents" going forward? Can't find this anywhere on their website despite them having known this for at least 14 days...
Topfi··on Jev in 25 Lines of Python
Uninformed hype for their startup. And they did a great job.
Topfi··on Jev-Leftpad
Nah, for ≥11 spaces we should fan out to a GPT-6 Astra agent. On light reasoning of course, lest we be wasteful.
Topfi··on Jev-Leftpad
Man-made horrors beyond my comprehension, neat.
Topfi··on Scott Jenson: Are we going to use the same Desktop UX forever? [video]
I am roundly miffed that I somehow missed that Scott Jenson was holding a talk in my very country. As always, great talk, his efforts in HIG are something every one of us benefits in ways large and small daily.
Topfi··on Pareto 26.9 Model Card
Very limited information, dare I say the least informative model card I've ever read. Then again, can you have a model card for what's seemingly a router?

Have to say that I like the naming scheme, just year and month over an arbitrary version number, I really vibe with that.

Topfi··on OpenJev
Unless I misunderstood what they wrote, I read parallelized in the diffusion sense, akin to GemmaDiffusion and Inception Labs models. Incidentally, Mercury 2.5 is truly groundbreaking, giving it a try is highly recommended.
Topfi··on OpenJev
Jev is, as far as I understand, essentially very optimised for zero shot classification [0]. Something like BERT could be and has been tuned to provide similar "decision making" at a similar latency and cost advantage quite some time back. Advantage over full on LLMs is mainly the efficiency and of something like Jev over e.g. the encoder/decoder based classifier I had in front of an LLM to route to different prompts depending on the users likely needs, that Jev does perform at a more consistent level, allegedly roughly akin to GPT-5.6 Terra, but at the lower cost and latency. Currently testing that, but seems promising, if Jev classifies at or above Terra level, I see no reason not to leverage it.

Can add that I tried using a heavily pruned mt0 based model for structured classification along with structured output for local tagging and simple renaming suggestions. While it does work, the balance is hard to get right for the machine I was targeting as a minimum spec (Macbook Neo), so that's on ice. Focusing on one of the tasks easily goes below 100mb with solid latency across all EU Latin script languages, but the second you add a few, it's simply not in the quality budget, so while LLMs can do anything Jev and similarly focused models can, it comes at a literal cost. Could maybe accomplish the goal with multiple models (BERT+mt0+...), but that get messy.

In general just happy to see a bit of the millions flooding into the industry being used to improve on less flashy but immensely useful solutions. It's amazing that you can technically use LLMs for most tasks, but not every org has a near infinite budget and there is still a lot to gain from applying more recent learnings to old solutions along with just updating their training data to the current year. Also makes business sense, competition on frontier or mid-tier LLMs is vicious, focusing on an underserved niche with clear application is clever.

[0] https://huggingface.co/tasks/zero-shot-classification

Topfi··on OpenJev
> If you have any comments about our WEB page, you can write us at the address shown above. However, due to the limited number of personnel in our corporate office, we are unable to provide a direct response.

A profoundly polite way to tell someone to stuff it.

Page 1 of 21Next →