15,843 karma · joined March 11, 2013
Website: https://kay.is
Blog: https://fllstck.dev
GitHub: https://github.com/kay-is
Wasn't worth it.
Not as fun to write as Python and Nim, but I don't have to write it.
Hrm.
Location: Germany
Remote: Yes
Willing to relocate: No
Roles: Technical Writer, Software Engineer
Homepage: https://kay.is
LinkedIn: https://www.linkedin.com/in/kay-plößerThey are well suited to check input for rule violations and are much cheaper to train and run than decoder models.
The downside is, they can't fix issues by themselves.
However, that would just change the weights values and not their dimensions.
All these rules would be way less cumbersome if they didn't come with a bunch of literal paperwork.
Why do the cache hit rates seem to vary so much between harnesses?
I use pi, which is very minimalist, and I get a hit rate of ~99%. Paying like $1 a day for Flash. Yet, the hit rate mentioned on OpenRouter is only ~79%.
I didn't do much agent coding and had a mix experience.
1. It would build something that was in the spirit of what I wanted, but unusable in practice.
2. It would build something quite useful, but only the public APIs were nice, the deeper code layers would get more and more convoluted.
3. It would built what I wanted and it would have okay-ish code.
However, for 3. I also had to add a custom AGENTS.md, many more code example, extra repos as subtrees, and review any code that had new concepts.
Much more work, but still much less than typing it all by hand.
So I'd guess it's API level.
Pro is ~50% more expensive than Flash.
Both need babysitting.
Plan, split in small tasks, give it docs, types, tests, linter, best practice examples, etc.
Always start a new session when starting a task.
Do regular manual sanity checks, and tell it to find issues in the codebase.
I pay like $1,50 per day for Pro.
What is the parento frontier?