Here's an updated eval with the proper models https://a3bmfqfom3.evvl.io/
13,210 karma · joined October 12, 2007
Twitter: http://twitter.com/mbuckbee
email: mike@expeditedsecurity.com
Here's an updated eval with the proper models https://a3bmfqfom3.evvl.io/
There's a lot of "tone" in it as she's not trying to anger these folks, but also it's quite serious, but also there's just everything else happening in medicine.
Feels like a great use.
They all did pretty well at a more "formal" tone, but GPT4.1 was the only one that didn't make me cringe with a "casual" tone.
[edit] fwiw, grok was also the fastest+cheapest model, claude was slowest and priciest.
I'm not particularly happy about that outcome as I wish we had more locally run AI models for reasons of privacy and efficiency, so this is more just a warning that at present there are some severe tradeoffs.
1 - https://sendcheckit.com/blog/ai-powered-subject-line-alterna...
1. "Mythical Man Month" which is the shorthand for a whole book + concept that you can't just throw more people at a software development project and get linear productivity improvements as the communications overhead (meetings, emails, mistakes due to poor assumptions, etc.) deeply eat into the raw number of productive hours that a new person added to the team brings.
2. AI automation tools (Claude Code) are often described as a "junior developer" which is an imperfect comparison as while you could potentially sort of set them up that way many people use them as more of a singular force multiplier.
I use them to work on many more projects in many more ways and ship far more than I could even if I had a "junior developer" sitting alongside of me as there's not the same level of communication needed.
OpenAI (or whoever) crashes and can't pay for the order leaving the memory makers in a tough spot.
1. Tests have always been both about the function of the application, but also the communication of what should be occurring to the larger team or yourself six months down the road.
With automated software development the communication with the LLM itself is a much larger part of it so I feel like it's "ok" to have lots of easy tests that are less about rigor and more about "yes this is how this should work"
2. Ideally we're going to get to the point where the tooling allows for adversarial agents with one writing code and one writing tests. Even for now just popping open a separate terminal window and generating+running tests in it from your main coding terminal is helpful.
There's still space for creativity, novelty, invention and human intuition.
I take your point, but think that it's also maybe too far.
That being said, I had enough issues with Bunny and CF debugging across regions that I made this free tool to do both remote HTTP and TCP traceroutes to keep my sanity: https://dnsisbeautiful.com/global-http-availability
That's not to say that the in browser isn't valuable for privacy+offline, just that the standard case currently is pretty rough.
https://sendcheckit.com/blog/ai-powered-subject-line-alterna...
I'm curious if there's a "good" way to do this.
What was surprising was that all the prefix+suffix variants of app/now, etc. were also taken so this was really just me trying to push it hard the other way.