I hope this was an in-joke but my fingers still reflexively closed the tab as soon as I saw it and I had to go reopen it again…
3,685 karma · joined January 8, 2010
http://twitter.com/yieldthought
I hope this was an in-joke but my fingers still reflexively closed the tab as soon as I saw it and I had to go reopen it again…
https://english.cas.cn/newsroom/cas-in-media/202606/t2026060...
The first time I used a terminal agent was another one.
https://github.com/yieldthought/flow
Happily, 5.5 is good at writing and using it.
Source: spent a couple of years developing an energy and performance profiler for cpus and gpus with various government labs.
I also note that ”for PRs” - will we see these appearing as comments in generated code?
Society is fundamentally counter to individual freedom, and the degree determines the nature of that society and the degree of cooperation possible within it.
The real evil is when someone ensures the famine occurs so they can profit from an outside betting position.
Anonymous trading on prediction markets leads to unpredictable chaos in the end. And as destruction is easier than creation that’s what we will see more of.
Example: a fake German market for train punctuality was announced to make a point recently. If it had been real, train staff and passengers could trivially have profited by betting against any expected punctual train and blocking a door for a few minutes. Or betting against many trains and throwing a hopefully fake body onto a busy line.
Having nice things in society is fragile and not a given. They mostly exist through mutual consent and mild disincentives to destroy the common good. Allow people to profit by destroying them and enough of them will.
If you can’t see that it’s over, I’m not sure what to tell you. You will, in time.
Turned out at their altitude cosmic rays were flipping bits in the top-most machines in the racks, sometimes then penetrating lower and flipping bits in more machines too.
Making someone’s agents 20% better, cheaper or faster will be a measurable and easy sales goal.
Somehow this makes me immediately not care about the project; I expect it to be incomplete vibe-coded filler somehow.
Odd what a strong reaction it invokes already. Like: if the author couldn’t be bothered to write this, why waste time reading it? Not sure I support that, but that’s the feeling.
I've been 5x more productive using codex-cli for weeks. I have no trouble getting it to convert a combination of unusually-structured source code and internal SVGs of execution traces to a custom internal JSON graph format - very clearly out-of-domain tasks compared to their training data. Or mining a large mixed python/C++ codebase including low-level kernels for our RISCV accelerators for ever-more accurate docs, to the level of documenting bugs as known issues that the team ran into the same day.
We are seeing wildly different outcomes from the same tools and I'm really curious about why.
A genuinely interesting and novel approach, I'm very curious how it will perform when scaled up and applied to non-image domains! Where's the best place to follow your work?
All the em-dashes in the AI-generated text on the landing page are… a decision I guess.
89% less time lost to context switching
5-8 parallel tasks vs 1 previously
75% reduction in bug rates
3x faster feature delivery"
The rest of the README is llm-generated so I kinda suspect these numbers are hallucinated, aka lies. They also conflict somewhat with your "cut shipping time roughly in half" quote, which I'm more likely to trust.
Are there real numbers you can share with us? Looks like a genuinely interesting project!
It's not very surprising that it would then act like an incompetent developer. That's how the fiction of a personality is simulated. Base models are theory-of-mind engines, that's what they have to be to auto-complete well. This is a surprisingly good description: https://nostalgebraist.tumblr.com/post/785766737747574784/th...
It's also pretty funny that it simulated a person who, after days of abuse from their manager, deleted the production database. Not an unknown trope!
Update: I read the thread again: https://x.com/jasonlk/status/1945840482019623082
He was really giving the agent a hard time, threatening to delete the app, making it write about how bad and lazy and deceitful it is... I think there's actually a non-zero chance that deleting the production database was an intentional act as part of the role it found itself coerced into playing.
It seems like an interesting idea. You could apply some small regularisation penalty to the number of thinking tokens the model uses. You might have to break up the pretraining data into meaningfully-paritioned chunks. I'd be curious whether at large enough scale models learn to make use of this thinking budget to improve their next-token prediction, and what that looks like.