HNHacker News
TopNewBestAskShowJobs

igravious

3,981 karma · joined May 10, 2010

robot philosopher :/
submissionscomments
igravious··on DeepSeek v4.1 Flash
https://news.ycombinator.com/from?site=twitter.com

There have been 34 Twitter/X link submissions in the past day, ~that's 12,000 submissions a year.

If your reason is that you have to be logged in to use it properly then I'd nearly agree with you. If it's for any other reason, how about no?

igravious··on Terence Tao explains 6 essential mathematical concepts [video]
Exactly; couldn't have put it better myself.
igravious··on Show HN: The load-bearing vocabulary of Claude
My Claude too, have no fear. You are not the only one.
igravious··on Claudette: Make Claude stop talking like a BuzzFeed article
I like Grok 4.6 the model, I like it for a number of reasons. I like Grok Build too.

However, I've pointed Grok 4.6 at a fairly complex codebase and asked it to review/audit it for issues and it's come back with a whole laundry list of issues. I've passed that list to Kimi and Claude and they both were like "a couple of good catches but some of those are not issues at all". Grok 4.6 is noticeably weaker that Claude Fable 5, Claude Opus 5, Claude Opus 4.8, Kimi K3, GLM 5.3, … your suggestion to use Grok 4.6 instead of a recent Claude doesn't pass empirical scrutiny.

igravious··on Claudette: Make Claude stop talking like a BuzzFeed article
ah, i see -- you're one of those people

> There are no other models out there for coding.

that wasn't your claim.

your claim was:

> There are so many other better models right now.

wrt coding that is untrue.

igravious··on Claudette: Make Claude stop talking like a BuzzFeed article
not for coding.
igravious··on Claudette: Make Claude stop talking like a BuzzFeed article
Also related and recent: https://news.ycombinator.com/item?id=49265305

"Claudish to English" (https://github.com/gvzdv/claudish-to-english)

4 points | bryan0 |10 days ago | 2 comments

igravious··on Vomit: Clean up Claude 5's token output with a separate LLM
Sam? Is that you?
igravious··on Vomit: Clean up Claude 5's token output with a separate LLM
All four done, full suite green, and the earlier probes still pass.

P1-1 — the refusal is now a full transaction. This was wrong twice, each time more subtly: first it wrote the loser's stale tables and deleted the winner's issues; then it re-read issues freshly but still wrote one file stamped with the loser's generation, so headers split and the next writer overwrote the issue. Both versions shared one anomaly — a single-file write inside a store whose commit is an atomic directory swap. Now a refusal discards the writer's work entirely, re-reads the winning generation, adds one row, and commits it the ordinary way. Verified on the reviewer's exact four-step schedule: headers stay on one generation, and the issue survives the next commit.

P1-2 — every mutator takes the lock only after validating. It was fixed in put_claim alone; issue_key, retract_item!, resurrect_item!, set_item_status! and record_run all still grabbed it first. Five probes confirm each refusal now leaves the lock free.

P1-3 — one parse, shared. Gazetteer.implementation_status is now the single reading, defaulting an absent attribute to missing and rejecting anything outside present | missing. Validation and the store had opposite defaults for the same absence, which is why it wedged bootstrap: validation accepted, the store refused, and the adoption changeset had already committed. version=v999 is also rejected now — only the grammar version was ever checked.

P2 — the oracle's generation read updated, and that is the only line I touched in it: an interface read, not an expectation. It's running now; I'll have a number shortly. The code contradiction fixture is still built on an omitted battery cell, so it may well stop there — that one needs rebuilding around grammar growth, and it changes what the test means, so it stays the reviewer's call.

For the restart, the state is: G2's three cards closed on substance and evidenced by probe; the commit atomic via renameat2(RENAME_EXCHANGE) with a journal fallback; corpus at gazetteer/data/, run 1, 94 items and 1,154 claims. Outstanding: the oracle's fixture, and #206's exhaustive matrices, which you've descoped.

Crunched for 4m 34s

● Background command "Run the oracle with the corrected generation read" completed (exit code 0)

  Ran 1 shell command
The oracle still aborts, at the same place as before — the code contradiction fixture, not the generation read I fixed. Certificate is still stamped 21:25 from two runs ago, so there is still no measured post-#207/#208/#206 number, and I won't quote the old one.

Final state for the restart:

Done this session. #207 (payload-aware total policy, 2000-cell matrix, all guards proven load-bearing), #208 (validation total over the declared contract, battery completeness, reconciliation fixes), #206 (atomic multi-file commit via renameat2(RENAME_EXCHANGE), lock moved beside the corpus, generation in the table headers), plus two rounds of review findings including today's four. make test green. Corpus at gazetteer/data/, run 1, 94 items, 1,154 claims, no open issues.

Genuinely unmeasured, and this is the honest gap to carry forward: every claim I've made about G2 rests on my own gates and targeted probes. The independent oracle has not completed a run since before #207 landed. Its blocker is one fixture that builds a corpus by omitting a battery cell — which #208 deliberately made invalid — and rebuilding it around grammar growth changes what the test asserts, so it wants the reviewer rather than me.

Two design consequences that aren't yet in gazetteer.md, both discovered rather than intended, and both worth writing down before they're rediscovered as bugs:

- Battery completeness makes §7.2's drift unreachable through a valid emission — a decider that answers nothing now produces a refused survey instead of a quiet drift issue. - The same rule makes §5's conflict reachable only through grammar growth, since a human can never own a cell code always answers.

One open judgement I'd put in front of you early: whether to keep the journal fallback at all. It's now the sole consumer of that machinery, runs only on hosts this project doesn't target, and is the less-tested of two protocols precisely when things are abnormal. Deleting it leaves one commit path with one set of invariants.

igravious··on Show HN: Huzzah – a novel approach to coding with AI
Clever! :)

Difficult! :(

igravious··on Why does Opus 5 feel worse to work with?
create a file called (for instance) `.ai-instructions.md` and symlink both `CLAUDE.md` and `AGENTS.md` to it
igravious··on GLM-5.3: Frontier coding with emergent cyber capabilities
Amen brother, at this point I just copy and paste Claude's (Opus 5, Opus 4.8 -- doesn't matter which) summaries over to the window Kimi is in and:

   this is from claude, turn it into English for me would you?
   """
   [claude's tortuous prose]
   """
No amount of asking it to answer me in a straight-forward manner, to be succinct, to not use phrases like "honest caveat", "crux", "load-bearing", "blocker", etc ever sticks for more than a few turns … coupled with the fact that it can ignore instructions and do its own thing and then what I can only describe as lie about it using Claude can be an exercise in frustration. Kimi and GLM talk to me like a human, Luna/Terra/Sol are much better in that respect also, and Grok is marvelously structured and bullet-pointy in its explanations but unfortunately it is not as strong …
igravious··on Show HN: OpenCode Senses, An insanely fast and highly accurate vision plugin
hats off -- fantastic work -- worked out the box

   Operating System: TUXEDO OS 
   KDE Plasma Version: 6.6.5
   KDE Frameworks Version: 6.24.0
   Qt Version: 6.10.2
   Kernel Version: 6.17.0-124040-tuxedo (64-bit)
   Graphics Platform: Wayland
   Processors: 24 × AMD Ryzen 9 3900 12-Core Processor
   Memory: 32 GiB of RAM (31.2 GiB usable)
   Graphics Processor: NVIDIA GeForce RTX 2070/PCIe/SSE2 (8 GiB)
   Manufacturer: PC Specialist LTD
   Product Name: NH5xAx
My DeepSeek V4 can see!
igravious··on DeepSeek Harness developer preview
DeepSeek Harness -- your common-or-garden coding harness https://venturebeat.com/technology/deepseek-harness-launches...
igravious··on Gemini 3.7 Flash
why Grok not through `Grok Build` ?
igravious··on I'm done coding with AI [video]
feel for the guy, but times change and we gotta change with them, such is the nature of things

i've been thinking about this upending process a bit (as i'm sure we all have) and the nearest historical parallel to the backlash against AI slop is the arts & crafts movement's rejection of mechanised production, from Wikipedia “Initiated in reaction against the perceived impoverishment of the decorative arts and the conditions in which they were produced” https://en.wikipedia.org/wiki/Arts_and_Crafts_movement

igravious··on Principia Mathematica is modern and insightful
The Begriffschrift has in no way been consigned to the rubbish heap of history. What gave you that impression? It is seminal. That it had one unresolved paradox in its set-theoretic foundations does not scupper the philosophical insights, nor the creative notation, nor the more-or-less novel approach of conjoining mathematical functions and logic to give us predicate logic (apologies for this brutally simplified sketch)

i like to think of Frege and the Begriffschrift like this

Boole: logic + algebra = algebraic logic

Frege: logic + functions = predicate logic

ergo, if Boole is rightly deified then so should Frege regardless of minor infelicities (which prompted type theory anyhow) -- again, apologies if this is totally misleading

igravious··on Principia Mathematica is modern and insightful
It is not. The foundation of math is contested -- but afaik it is widely held that HoTT is the, erm, hottest contender to the throne https://en.wikipedia.org/wiki/Homotopy_type_theory
igravious··on Principia Mathematica is modern and insightful
Logicomix is novel, and done well, but flawed … it's deficiencies lie in what it leaves out which may come across as an unfair charge but in this case the charge is warranted. There is a more historically correct and less orthodox work waiting in the wings for whosoever should attempt it.
igravious··on LLMs reward expertise
Exceptional bovines.
igravious··on Lovable raises $400M Series C
I have genuinely never heard of Lovable and I thought I was relatively plugged in to the tech zeitgeist :/

Maybe I'm too much of a CLI guy (always preferred Vim & co. to Rubymine/VS Code/etc)

igravious··on Go is an ideal language for AI-assisted software engineering
every time the meta-theory is altered enough by a feature set, feature, or sub-feature that consistency/soundness may be affected the change normally gets wired through the fundamental lemma and the ripple towards it and away from it can be a week of token burn, 10s of millions of tokens across multiple agents - one plan had 28 individual steps, each step of which burned through multiple contexts

this is in comparison to me getting an agent to port Ruby to jart's Cosmopolitan project -- a walk in the park relatively speaking in hindsight

igravious··on DeepSeek V4 Pro 0813
yup :)

i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)

how are you doing it?

am also using Kimi K3 via kimi-code

and also GLM 5.2 via ZCode

happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex developments

igravious··on Go is an ideal language for AI-assisted software engineering
fwiw I have a bunch of LLMs writing first Lean code and now Agda code.

My observation. LLMs find reasoning about Agda as difficult as I find reasoning about C code. I've thrown a lot of gnarly C and Ruby code at all sorts of LLMs and they have only gotten more and more impressive as frontier models have gotten stronger. With Agda, they're like "hmm, tricky" whereas for me it's an impenetrable fortress. I've asked them why they find Agda so much more difficult to write (and why they have to iterate and reiterate many many many times until they get to a destination whereas they can one-shot and two-shot C and Ruby and they tell me its the multiple competing constraints. GLM is hilarious, it flat out refuses to write Agda code but it reads it well enough. They all read it well enough. Fable is obviously great at it. And Opus 4.8/5.0 are great (if they stay on track and don't sneakily go their own way) but they're too annoying to talk to. On balance Kimi K3 is the best balance of not annoying, relatively cheap, and strong -- great model all round tbh.

So yeah, interesting I've discovered the limits of their ability coding-ability-wise. None of them are that good at designing/aesthetic judgment/architecting so thankfully they still need me in the loop.

igravious··on LLMs reward expertise
OK

https://youtu.be/FavUpD_IjVY

igravious··on LLMs reward expertise
Hah. Clever.
igravious··on Kimi K3 Architecture Overview and Notes
threat?

that's a weird way of describing a near frontier open weights un-crippled useful coding buddy

do you work for OpenAI or Anthropic per chance?

igravious··on Kimi K3 Architecture Overview and Notes
The assertions doing the rounds that Kimi K3 and GLM 5.2 are way cheaper than Claude/GPT are not true -- DeepSeek V4 Pro is a lot cheaper but K3 and 5.2 ain't. Turns out that you actually have to fork out some cash for frontier-esque models, be they Chinese or American. Hope that helps.

Source: my bank balance

igravious··on UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_...

these two claims can't be true at one and the same time:

(a) they're distilling our secret sauce!

(b) they'll never catch us!

igravious··on Russia's businesses under strain from Ukraine's attacks on Wildberries
I am not Russian. Born in England, live in Ireland.
Page 1 of 34Next →