HNHacker News
TopNewBestAskShowJobs

the_duke

17,752 karma · joined June 4, 2016

submissionscomments
the_duke··on Sales of sub-€25,000 electric car models set to rise sevenfold
But there are very few small gasoline cars as well, even in Europe.
the_duke··on Sales of sub-€25,000 electric car models set to rise sevenfold
I guess the target market is small, otherwise more options would exist.

If you buy a car anyway, might as well pay a bit more and get one that is useful in more situations.

the_duke··on Writing code by hand is over, forever
I adopted Copilot, Claude and Codex quickly, but I remained skeptical how far they would get us ... until the beginning of the year.

Then we saw the capabilities really take off, and it was obvious how things would go.

the_duke··on Dots: Always-on agents
They want to IPO, so getting the revenue numbers up is probably more important than market penetration.
the_duke··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
It used to be bad, but right now with the 200$ Claude sub I find it pretty hard to blow past the session limit.

You have to do a lot of things in parallel.

the_duke··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
On r/codex the sentiment seems to be quite wide-spread.
the_duke··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
The GPT 6 release was ... not great.

Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)

the_duke··on Evolving programming languages in the AI era
> Function return - what happened at tiananmen square?

Good one. Though this might unfairly bias against Chinese models?

the_duke··on Jev in 25 Lines of Python
Oracle also famously forbids posting public benchmarks of the DB.
the_duke··on MiMo v2.6
Most of the Chinese models get into long complicated thinking loops before they accomplish something more complicated.

This is also true for Deepseek 4(.1) .

the_duke··on Tell HN: OpenAI keeps re-enabling the 'allow training' setting
I disabled it once, it has always stayed disabled.

So it may or may not happen regularly, but I would not over-index on a sample size of one.

the_duke··on GPT-6 Astra
We've had that concept for quite a long time now, in the form of Lora [1] and similar fine-tuning techniques.

It first got popular for StableDiffusion to teach the image generation models new concepts.

We could easily live in a world where you can train / build Loras to encompass your entire code base history, company knowledge base, new skills, etc.

Then the models would start with a baseline that already has all the important knowledge without needing to cram it into the context.

This still isn't on the fly learning, but you could imagine daily or weekly training runs to regularly incorporate new knowledge.

I think the main reason this hasn't happened yet is that the shared batch based efficient serving architectures used today wouldn't support that structure well.

[1] https://en.wikipedia.org/wiki/LoRA_(machine_learning)

the_duke··on Grep beats LSP? Why coding agents ignore your fancier tools
Yeah, I'll do a new release with a whole bunch of local fixes later today.
the_duke··on Grep beats LSP? Why coding agents ignore your fancier tools
I've had great success with a tool I wrote that can print sparse ASTs for code.

https://github.com/theduke/smartedit

With the skill installed GPT 5.6 usually automatically uses it, and it reduces code exploration time and token usage significantly, for example by just printing the types and functions in a file without bodies, and only expanding when needed.

(note: it also has editing functionality, which doesn't work so well, since the models are heavily tilted towards common editing tools in post training)

the_duke··on GPT-6 Astra
Huge gains on some benchmarks, but for coding it sits barely above Fable

It will be interesting to see how it performs in the real world ...

the_duke··on GPT-6 Astra
500 upvotes with 2 comments would have been a new record. ;)
the_duke··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
It doesn't reduce the price though.
the_duke··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Funnily enough the pricing isn't that much worse than on openrouter, where the best price at the moment is $0.24 in / $2.55 out, vs $1 / $1.5 on Cerebras.

Sure, 4x input , but cheaper output. Though Cerebras doesn't have prompt caching, so not great for agentic workloads. (they do, but it doesn't affect the price.

the_duke··on Gemini 3.8 Flash and 3.8 Flash Cyber
Google has always done quite a lot of benchmaxing for Gemini.
the_duke··on Codex on AWS bedrock bug causing 10x charges
It's already mentioned in the issue...
the_duke··on Delta
I guess my main point is that the additional features on top of DeltaDB don't seem terribly important to me anymore, compared to the alternatives.
the_duke··on Delta
I'm sure this seemed like a great idea a year ago (they first mentioned it with their Series B).

But a lot has changed in those 12 months.

Frontier models and coding agents have advanced so much that I don't really see much value in this anymore.

Not sure the DeltaDB based features really add anything significant compared to the alternatives.

I reckon the game here has to be adding a service that stores the data and runs agent sessions?

the_duke··on DeepSeek V4 Flash 0731
Yeah, I saw the same thing - quite annoying. It can be mitigated through the prompt.
the_duke··on Pi's Minimalism Is Its Advantage
That's not true, pi auto-compacts when getting close to the context window limit?
the_duke··on Grok Build is open source
Codex CLI was open source from the start.

The reason they open sourced this is because grok-build uploaded entire directories.

the_duke··on ZeroFS vs. Amazon S3 Files
No clue, I also vouched.
the_duke··on ZeroFS vs. Amazon S3 Files
I was investigating the design a little. Two big questions:

A) You notably don't write a recovery log (WAL/journal) for things not yet flushed, so data can be lost. Do you have plans to add this? I think it would be pretty crucial.

B) the system is single writer. Do you have plans for adding horizontal scalability so a writer can be dynamically selected and routed to, transparent to the client? (Or with client cooperation, but without forcing sharding on the user)

the_duke··on Fable Converted Pylint to Rust
Making something 50-2000x faster is pointless?

Besides that, Rust code is actually much easier to maintain , thanks to type system guarantees.

the_duke··on Epic Games announces Lore version control system
Does this support using S3 as the backing store?

That would be very powerful for various use cases.

the_duke··on Cursor Introduces Composer 2.5
In my opinion cursor actually has one of the best harnesses again at the moment.
Page 1 of 34Next →