HNHacker News
TopNewBestAskShowJobs

rbuccigrossi

11 karma · joined July 15, 2020

submissionscomments
rbuccigrossi··on ARC-AGI Leaderboard
No, I believe you are misunderstanding the quote. Each “question” in ARC-AGI-3 is a game that has hidden rules that you can understand if you look at the game board long enough. This quote means that Opus 5 is looking at the game board, figuring out the rules, and writing out the rules before it makes a single move. You can do the same thing if you go to the ARC-AGI-3 website and try some of the games.
rbuccigrossi··on Is AI Progress Real? Four Independent Metrics Show It
The others are an IQ test (TrackingAI), a test of Ph.D. questions across multiple domains (Humanity's Last Exam), and graphical pattern matching (ARC-AGI-2).

What's interesting is that while they are rather different in nature (yes it is odd that METR measures clock time as opposed to iterations etc.) but the behavior of the resulting improvement curves are extremely close.

That's the punchline: 4 independent measures point to the same conclusion. That increases the chance that the conclusion is correct.

rbuccigrossi··on Is AI Progress Real? Four Independent Metrics Show It
The text is my transcript of the video formatted by Claude and with sources added at the end.

I do use deep research (across Gemini, ChatGPT, and Claude) to gather background and ideas, and Claude for editing. I started machine learning and computer vision research back in 1995, studied NLP, and dove into LLMs with GPT-2, so I wouldn't be surprised if my writing has been deeply influenced by AI.

rbuccigrossi··on Is AI Progress Real? Four Independent Metrics Show It
Author here. In short there are 4 different metrics:

- METR's time horizon

- TrackingAI's offline cognitive test

- Humanity's Last Exam, and

- ARC-AGI-2

that have lasted longer than 2 years (though ARC-AGI-2 is now saturated).

When plotted in the linear domain, they all have an exponential (hockey-shaped) curve, but the interesting thing is that the bend happens right at Q4 of 2025 (right when Gemini 3, Opus 4.5, GPT 5.2 all come out).

rbuccigrossi··on Local AI Hardware: Break Even in 2.6 Years?
In short, running a $3,299 GMKtek EVO-X2 (Ryzen AI Max+395 with 198 GB) 24/7 with the Gemma 4 26B-A4B model, being as generous as possible, only saves you $1,279.07/year in inference costs. (120 t/s for output tokens at $0.34/M.)

So, that's how you get a break even at 2.6 years...

rbuccigrossi··on Show HN: Decoding the Language Machine – AI video series and CC repo
You're first :D
rbuccigrossi··on Threatening AI Does Not Make It More Useful. Why Sergey Brin Is Wrong
We work in the arena of automated AI workflows where consistency of success is vital. When you threaten an LLM you are drawing the LLM into the texts where threats occur (flame wars, parody, etc.). So intuitively you would expect it to work sometimes, but also fail with even more ardent refusal (increasing the variance of success).

Jailbreak approaches like "Bad Likert Judge" ( https://unit42.paloaltonetworks.com/multi-turn-technique-jai... ) and similar persuasive techniques (see https://xthemadgenius.medium.com/how-persuasion-techniques-c... ) move the text domain to more policy, analysis, or scientific papers, where deeper analysis, discussion, and compliance is the norm.

So I'm curious about the extremes (variance) of success with threatening vs. polite discussion, but I haven't seen direct research on that.

rbuccigrossi··on Threatening AI Does Not Make It More Useful. Why Sergey Brin Is Wrong
Treating an LLM with respect is not about pretending it has feelings; it’s about understanding that every word in your prompt is a signal that shifts the probabilistic landscape from which the model draws its answer. It’s about probability, not personality.