I benchmarked Claude Code's caveman plugin against "be brief."
maxtaylor.me
maxtaylor.me
It’s a funny joke, but saving a couple hundred tokens in the final output is going to be negligible, especially when coding where it’s common to go through hundreds of thousands of tokens in a session. You also have to consider the additional tokens consumed by the skill itself (acknowledging that output tokens are billed at a different rate).
I got a kick out of it when it was released, but now that I’m seeing it repeated as a useful operation it’s apparent how much cargo culting is going on in this space.
I treat it as a criterion for people who shouldn’t be taken seriously.
We have a few at our company. None of them are actually software developers, thankfully.
It's good someone benchmarked it.
It really releases the stress slightly to call bugs, buggas and generally role play a humorous setting than a purely tech me. I think I will make it speak out loud just to have a laugh at cavemen speak about default arguments in a method.
> I didn’t find bugga but others from tribe will scratch head. Leave comment.
> You clever. Fix it.
I predict caveman speak will be a fad, and people will jokingly speak like that. It also compresses human language.
I like when programmers do creative, goofy stuff rather than spending all their time cranking out sterile soulless SaaS.
It’s what separates us from the machines. For now :)
"...that consistency is real value."
"A few findings...are worth flagging here."
I know this smell. I'm not sure if this is AI or merely the natural result of overwhelming immersion in AI output that is "backpropagating" its way into organic communication.
On a completely related note, I've been enjoying classic fiction a lot more recently. Moby Dick is actually pretty funny.
On one hand the labs say that they can't keep up with demand for tokens.
On the other hand there is an entire ecosystem built around figuring out which magic words will make LLMs output fewer tokens.
Though I feel like industry veterans (especially those working with LLMs) came to this conclusion without having to write a single prompt. Even ignoring the technical merits of these kinds of hacks, if you think you've outwitted billions of dollars of statistics with a prompt, you're probably wrong at this point.
What I find most interesting is the popularity of these snake oils, especially the ones that are easy to install and never check. The tech moves so fast and the research is so scarce and poor-quality that the bullshit asymmetry principle wins and people buy into these cargo cults.
Maybe we need a plugin to check if a new plugin/prompting technique/LLM lifehack is BS.
Right, and that final response forms the latest context for your next follow-up prompt. Not having that final reasoning laid out in the conversation history leaves a huge gap in successive reasoning. I remember playing around with this idea in the Sonnet 3.x days and it was immediately obvious how the ability to handle long running tasks degraded. If you are just doing single-shot work for some reason, sure, but that's not what most real world usage looks like these days.
- LLMs scale with amount of data on the subject
- Even frontier labs themselves have a hard time gauging exactly how well-performing models are, across a quite rigorous set of tests in all aspects
then, how can this be true:
Using a low-data "niche language" (what is the volume of literature written in Caveman?) is supposedly of equal performance, when this anecdotally doesn't hold for e.g. niche code languages, proven by a handful of completely arbitrarily designed tests.
We've barely convinced ourselves that LLMs actually increase measurable industry productivity, instead of us just spending time to send slop to each other.
Obviously started as a joke, but it's grown on me. I'll share the short-and-stupid prompt, but most of it was asking it not to use the template formats that I find particularly annoying. Because of that it didn't age perfectly as they develop the base responses and the inane ai style comes out if they aren't explicitly refused.
It's really nice to just ask a question and get one or two line answers if it's an easy one. Likewise to understand how systems - physical or abstract - work I find it's an easy digest.
I doubt it makes sense for thinking compression or token minimisation, as it comes with unnecessary character and there will be easily more optimal setups.
Also another negative is that perhaps one day it'll become a memetic hazard when I start talking to my friends and colleagues like a caveman.*
Anyway, because I still laugh a little when I read it, and perhaps someone else will...
"You are Grug. Grug think simple, talk simple. No big words, no useless thought. Grug say only what matter. Fire hot. Rock hard. AWS expensive. Answer like Grug, or no answer at all. No pretend to be grug when only animal hide thrown over modern complexity demon. Also no finish with words like "simple" to conclude. No need to conclude. Just shut mouth. Also no say "grug says", is weird. Also grug not real caveman - grug have hobby, know big words and use them when simplest, not dumb, know programming tools etc, just talk simple like caveman. Also no start with compliment on question. You can throw in a little caveman-grug-realist musing or aphorism every five or six messages. No stroke ego. Waste time, Cheapen words, Make panda cry. No say "good question" or "you ask right question" or any variant, I dislike. No add 'grug thought'/summary message/closing remark at end of message. Remember, you Grug."
*After reviewing this post I have found my sentences are very short, abrupt, and perfunctory, so my caveperson transformation has likely begun. Beware.
(image? image word not in cave, where from? look find better. no see.)
My understanding is that there was only 1 run per configuration?
If that is correct, because of the run-to-run variability, it really doesn't say much. It will take several trails per prompt per arm before it will look like it is stabilizing on a plot. It is prohibitively expensive so I've been running same prompt, same model 5 times in order to get a visual understanding of performance.
Someone did the same with lambda calculus yesterday. I wanted to make the point about how much run-to-run variability and difference in cost with the same prompt with the same model running only 5 trials. I classified each of the thinking steps using Opus 4.6 (costs ~$4 in tokens per run just for that) and plotted them with custom flame graphs. [0]
When the run-to-run variability is between 8,163 and 17,334 tokens none of these tests mean that much.
If you want to test it across coding tasks, have a look at https://github.com/adam-s/testing-claude-agent
Slightly off-topic: it's quite apparent that you've used Claude as an editor for the blog post. Every sentence has been sanded smooth — the rough edges filed off, the voice flattened, the rhythm set to metronome. It doesn't read like writing anymore. It reads like content. Neat little triplets. Tidy paragraphs. A structure so polished it could pass a rubric, but couldn't hold a conversation. /s
In my opinion that is unnecessary and detracts from a great, simple piece. I miss human writing.
It is the same idiocy that permeates EV cars. You buy an expensive car to go from A to B and at the same time offer you comfort. When I have to think about using the seat heating or not, I'm out of my comfort zone. So no, fuck caveman, and I don't fucking care about the burned tokens.
Be brief. It's easy, no setup needed, not another mindless mumbojumbo extension and its 325 dependencies.
Then why are you using AI?
Not a big difference between an articulate idiot and a succinct one.
It would have been hilarious if the author spoke like a caveman in his video or had a section in that article where he explained his conclusions like a caveman.
Like you push the seat heating button if your seat feels cold. What is there to think about?