17,752 karma · joined June 4, 2016
If you buy a car anyway, might as well pay a bit more and get one that is useful in more situations.
Then we saw the capabilities really take off, and it was obvious how things would go.
You have to do a lot of things in parallel.
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
Good one. Though this might unfairly bias against Chinese models?
This is also true for Deepseek 4(.1) .
So it may or may not happen regularly, but I would not over-index on a sample size of one.
It first got popular for StableDiffusion to teach the image generation models new concepts.
We could easily live in a world where you can train / build Loras to encompass your entire code base history, company knowledge base, new skills, etc.
Then the models would start with a baseline that already has all the important knowledge without needing to cram it into the context.
This still isn't on the fly learning, but you could imagine daily or weekly training runs to regularly incorporate new knowledge.
I think the main reason this hasn't happened yet is that the shared batch based efficient serving architectures used today wouldn't support that structure well.
https://github.com/theduke/smartedit
With the skill installed GPT 5.6 usually automatically uses it, and it reduces code exploration time and token usage significantly, for example by just printing the types and functions in a file without bodies, and only expanding when needed.
(note: it also has editing functionality, which doesn't work so well, since the models are heavily tilted towards common editing tools in post training)
It will be interesting to see how it performs in the real world ...
Sure, 4x input , but cheaper output. Though Cerebras doesn't have prompt caching, so not great for agentic workloads. (they do, but it doesn't affect the price.
But a lot has changed in those 12 months.
Frontier models and coding agents have advanced so much that I don't really see much value in this anymore.
Not sure the DeltaDB based features really add anything significant compared to the alternatives.
I reckon the game here has to be adding a service that stores the data and runs agent sessions?
The reason they open sourced this is because grok-build uploaded entire directories.
A) You notably don't write a recovery log (WAL/journal) for things not yet flushed, so data can be lost. Do you have plans to add this? I think it would be pretty crucial.
B) the system is single writer. Do you have plans for adding horizontal scalability so a writer can be dynamically selected and routed to, transparent to the client? (Or with client cooperation, but without forcing sharding on the user)
Besides that, Rust code is actually much easier to maintain , thanks to type system guarantees.
That would be very powerful for various use cases.