Tokenmaxxing is dead, long live tokenmaxxing
12gramsofcarbon.com
12gramsofcarbon.com
However, I think finding security vulnerabilities is one use case where it doesn’t matter. Tokenmaxxing is absolutely effective for that. We as an industry are in the middle of adopting very expensive, complex continuous fuzzers.
wow! That sounds like an unbelievable grift. Who were they such that anyone could possibly think that's a worthwhile investment?
Afaik messing with the context also pretty reliably degrades performance still. The model responses reference things that no longer exist to it and it becomes more chaotic.
The real usefulness of parallel or sub-agents is not that they run at the same time, its that they isolate noisy or self-contained context away from the main window.
Have you measured them beyond loc?
How many people were in the class?
I had a small training company, shuttered during COVID, and I used to charge 5K per day, for a group of up to 12 people. 5 days training = 25K. This is double, wow.
I would love to get back into training but getting enough volume to live and support a family on has been a challenge.
Personally, I get huge mileage out of LLMs, and yes, I care deeply about code quality, readability, and debuggability.
I've seen juniors absolutely rock with them.
And I've seen the exact opposite, where they just struggle to get good results.
In the end, I think the divide comes down to management experience. The people thriving are the ones who have led teams, especially teams of contractors, which is the best analogy for how you have to interact with an LLM.
Those folks know how to break down problems, provide the right context, and scope a task just enough to see the "contractor" succeed before letting them move forward.
On the other hand, individual contributors who are used to just grinding solo often struggle. They expect a one-shot miracle. They say, "Hey, my code is buggy, fix it." When the LLM inevitably hallucinates or steers them wrong, they give up. The results are completely different based on how you treat the tool.
They might just have a high quality of control and standards that it is hard to find that pattern with the LLMs.
I think fierce individual contributors are a lot more valuable in the era of llms as well. We as humans typically achieve better balance with new stuff when we allow backlash from new processes that start to trample on old ones without understanding AKA the Chester's fence.
Anyways, more of a ramble than my two cents.
In Hawaii I assume?
Like ... pivoting to the "metaverse" and changing the company name to show he's serious.
I think they're just a few decades too early.
Similar to cloud gaming, although that's much closer on the horizon.
Definitely not some measured, long term, rational out of the gate.
The goal isn't to have people work at converting wood into sawdust, the point is that if you wanna see if the tools are working you wanna see proof they're actually being used.
I'm sure there were some people cargo-culting this stuff, but suggesting that the people who run FAANG don't understand the dangers of bad metrics is... interesting.
(Of course, we've all had bosses that went to some marketing seminar and come back having been tricked^Wsold into buying some wizz-bang widget that we need to now integrate because of a sunk-cost fallacy, but I thought everyone was on the same page that this is not how normal procurement was supposed to work.)
> the point is that if you wanna see if the tools are working you wanna see proof they're actually being used.
That is way too charitable, people were being fired based on these metrics and people were absolutely talking about token burn as being a metric for productivity (do I really need to link the Jensen Huang quote?). That isn't an indication of this hysteria being based on "just trying to see if the tools work".
If you want to see if the tools work, why don't you just ask your employees? Like any normal employer would?
Though in theory power tools are faster than hand tools.
> Don’t count electricity and sawdust
I agree that it seems wasteful, but is there some better way to accomplish it at the scale of hundreds, or hundreds of thousands, etc? I'm personally doubtful that stubborn employees would switch even if a video provided internal metrics, videos, etc.
Awesome, can you please share those results? Surely they would be all over Nature, Science, IEEE, etc.
We couldn't build any of this stuff (rockets, LLMs, heart medicine) if the foundation was ill defined.
I think it's the second time I run into you like this, Temporal. I wish HN had a way to classify you as an "AI booster" or equivalent.
Yes. There's a lot of interesting things science has to say about water, very specific claims that took a lot of effort to discover, precisely formulate, and reproduce.
We're not talking about those. The whole LLM discussion on HN, as well as in the wider industry, is still stuck at the state where a large (or vocal) group of people refuses to believe water is wet. Yes, there is a similar group that tries to sell water as miracle cure, I'm not denying it - IMO both perspectives are dumb and entirely detached from obvious observational evidence that you can collect for ~free at home in 15 minutes. Example will follow.
There exist the equivalent of foundational, detailed studies on LLMs, at every level of rigor imaginable (with a caveat, it's hard to rigorously prove anything useful in software engineering; it's still largely opinion-driven field). But they're not part of the overall "AI hypers/haters" dynamics.
> I wish HN had a way to classify you as an "AI booster" or equivalent.
You can take any of the LLMs and have it vibecode you a user script in under 5 minutes, than you then can paste into Greasemonkey/Tampermonkey, and voilà, you have me labeled as "AI booster" or filtered out.
In fact, let me help you, I'll time it. I opened chatgpt.com in incognito (to emulate being a rando free user), and put the following prompt in:
> I need a user script I can paste into Tampermonkey on my Firefox that will clearly label user named TeMPOraL with robot emoji and some silly emoji, so I never forget when reading their HackerNews comments that they're an unapologetic AI booster.
Got back this script in under 10 seconds: https://pastebin.com/akEchvHd. Tested it, works out of the box.
This is the promised empirical example. It doesn't prove everything, but it proves something, and it took, end-to-end, a total of 1 minute to perform just now. You can collect many such examples over a single day by just trying. People who keep saying AI is useless and a fad and can't do anything useful, obviously never bother with even that.
FYI: I'm not an AI booster. I like AI, and I find it useful, but I'm not going out of my way to boost it. I just enjoy this topic, but more importantly - and I remain consistent in this - I point out bullshit that doesn't agree with obvious observable reality.
EDIT: try the example yourself, and post whether it works for you too - if it does, it's technically a peer-reviewed, replicated study, but I doubt it'll convince any of the naysayers of anything.
EDIT2: I have plenty of negative things to say about LLM capabilities and how irresponsibly people use them, and I do occasionally write about this (mostly at work, these days), but most HN threads on AI are not on this level - not anymore. They used to be more reasonable back in GPT-4 days.
They're not "3-4 trillion dollars in investments over 5 years" useful, nor "crammed into the throat of every employee on the planet, regardless of their actual job" useful.
The way they are pushed right now will lead to a very hard crash and probably lots of suffering. Also, you need a more advanced prompt for Firefox on Android :-p
Why not? They're a general-purpose technology, in the same category as "software" or "electricity".
> nor "crammed into the throat of every employee on the planet, regardless of their actual job" useful
They're potentially useful for anything that can be fed into computers (VLMs lifted the "that can be expressed as text" limitation, visual and audio tokens are not a separate category to text tokens anymore). That touches every single job people do in some aspects. Even though LLMs can't do physical work for people, they're still able to help with directing it and teaching it.
"Cramming into the throat of every employee on the planet" was already covered by many comments here, and the article itself - it's about forcing the obstinate holdouts to at least try.
> Also, you need a more advanced prompt for Firefox on Android :-p
No I don't; literally copy-pasted it to Tampermonkey on my Firefox on Android just now, and it works there out-of-the-box too.
Regarding LLMs, they are pushed too hard and too abusively by business people. Employees are being laid off and replaced with chatbots that don't do the job. Frustrating if support for McDonald's, risky if health insurance support. Also the financials don't make sense. AI companies are money pits. Money is ultimately production. We make X amount of stuff yearly, globally. We can't afford to through away 5% of X yearly on technologies that will probably have a proper return in 5 or 10 years. When we mis-allocate resources on scales like these, people die. Look at Communist centralized planning. For $3-4 trillion we could have solved a LOT of actual global problems.
LLMs are fine but they should have matured in the software dev domain for 2-3 more years and then non tech products would have followed.
After first failure (Gemini 3.5 Flash + NotebookLM), I run the other two (Opus 4.8 on Extra; GPT 5.5 on High) in parallel, and looking at their thought streams, I gave up and dug up the manual and read half of it, before the LLMs finished coming up with - wait for it - wrong answers.
Super frustrating. Doubly so, given that I use them for comparable tasks pretty often and they usually sail through them flawlessly. But this experience happens every now and then. It's only fair to report it, if only so you don't think I'm just AI boosting all the time.
Was the correct answer not "use the compartment marked with two parallel vertical lines"?
[ 2 | 3 ]
----------
[ 1 ]
1 is for powder detergent, 2 is for liquid detergent, 3 is for softeners and such.All three LLMs (Gemini twice, since NotebookLM) insisted I should put the powder detergent into leftmost compartment (2 on the ASCII diagram above). They referred to it by different numbers, but all gave some convincing justification why to put the detergent there. That's despite me posting photos of the compartment drawer, with symbols clearly visible. That's despite demanding they find the manual and cross-ref. I even asked two (Gemini and Claude) to label the actual compartment on the photos I took[0], and both produced some nonsense, with labels in all the wrong places. And they all insisted they're right and issued plenty of warnings about making sure I get this right or else bad things will happen.
BTW. I ended up posting a screenshot of the diagram in the manual to Claude with a passive aggressive comment. Looking at its "thinking summary" and tool calls now, I think at least Claude didn't process the image correctly and only saw:
[ 2 | 3 ]
------------
as those parts are blue, while the bottom is just in the same color as the entire body/frame of the machine. Maybe the contrast was too low for the models. But it was okay for humans, so it's not excusing much, especially that they all claimed to have found the manual, which had a high-contrast diagram.(Current experience tells me they probably didn't really check the diagram. I noticed recently that all major models seem to have gotten lazy when it comes to reading sources, and are also more than happy to lie about it.)
--
[0] - A method I often use with Claude when I'm not sure if it's dealing with spatial tasks correctly - I have it produce intermediate artifacts that involve modifying "ground truth" inputs - e.g. placing two map pictures on top to verify it solved the coordinate transform, and/or (like here) drawing labels and boxes on top of original photos. I found such requests to be helpful enough I set it as general rule for Claude now.
[1] - Which normally they'd spot, but for some reasons, they didn't.
> I have it produce intermediate artifacts that involve modifying "ground truth" inputs - e.g. placing two map pictures on top to verify it solved the coordinate transform, and/or (like here) drawing labels and boxes on top of original photos.
Like you I use intermediate artifacts all the time, but have never tried with visual elements. But how do you get Claude to modify images? Do you get it to output things to a canvas? html? use an external library?
At the risk of abusing the analogy further: many people aren't refusing to believe it's wet, they're observing that sulfuric acid is also "wet" and can look similar upon visual inspection, and there's a lot of harm coming along with the demonstrated capabilities, in addition to those capabilities themselves being fickle and inconsistent (not a desirable property for a good technology).
This isn't a problem of "doesn't know what AI can do"; yes, some people are misinformed, but you shouldn't dismiss all refusal to use AI as being misinformed. This is a problem of "knows what AI can do, and based on that informed position thinks it's terrible and should have careful guardrails around it".
https://www.nature.com/articles/s41567-026-03299-z
Scientists find more or less everything very interesting, even (especially?) things that are supposedly self-evident. You can both make a big splash disproving self-evident things, and much can be learned from it.
https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways
Or these?
https://www.forbes.com/councils/forbestechcouncil/2026/03/16...
Or these?
https://poll.qu.edu/poll-release?releaseid=3955
Yes, "results are in". They're all over the map, about productivity, about stress and churn, about trust, about public sentiment, etc.
But sure, if you want to tell people their productivity will be measured by token usage, they will certainly respond to that incentive by setting your checkbook on fire while they work on a job search.
If a company wants to provide AI accounts for people, along with guidance for usage and non-usage, that might well make sense for some jobs. It certainly makes sense for some uses. If they start measuring token usage, that's even worse than when companies tried to measure lines of code written.
because that would require actually admitting that employees are the people in an organisation who are responsible for the success of that organisation, rather than the people higher up the org chart.
maybe, just maybe, it would have been a better idea to engage with employees first rather than posting on linkedin about how everyone is going to lose their jobs.
cos it's the kinds of people trying to force this stuff on employees that are the ones who have been shouting about that from the rooftops.
Seriously, some of the most deranged things I've ever read were by relatively normal people trying to promote themselves on LinkedIn.
What people SAY does not matter nearly as much as what everyone KNOWS and it's pretty damn clear that AI is never going to be able to replace humans in complex domains. Every time a frontier lab announces a breakthrough it's pretty obvious that the setup was more complicated than "hey chat prove the Riemann hypothesis."
The world is gonna need skilled human beings to drive LLMs, no matter how desperately some people like to pretend otherwise.
However it takes some taste in engineering and perhaps some mathematical sophistication to figure these things out. “Just use AI,” is not a very convincing argument either.
It’ll take time to sort out, I wager.
Are you suggesting that changes to new production technologies are always driven bottom up by line workers? I'm guessing that historically that's rare.
But to give you an example, also roughly 0 companies made developers use Linux and still many developers choose it, so bottom up improvements happen in a decent chunk of cases. Nobody paid for PostgreSQL promotion. Or Python, etc.
It does, but for better or worse, it's an anomaly. Even now, maybe nobody was paid for PostgreSQL or Python promotion, but modern OSS tools and programming languages usually have a business backing it. Linux, too, wasn't commercially promoted until it was; RedHat isn't exactly a charity after all.
Conversely, no one paid for initial AI promotion either - ChatGPT exploded organically after release, and for the first year or two, companies had a problem because a good chunk of their staff, including especially non-engineers, discovered just how useful it was and wanted to use it at work, casually violating every internal policy, bylaw and even regulatory policies about data sharing. The massive spend on promotion - including first-party spend - came later, but at that point it was already obvious ~everyone is going to be buying it.
I suppose bottom-up vs. top-down may be in part about how mature a technology and industry is.
For one, software tools are cheap, especially with OSS in the mix. You're buying one "tool" and paying for operational expenses that scale with total usage across all company.
But secondly, and more importantly, the "consulting" and discussing was done over the period of last 3 years, by ~1 year ago the high-level conclusions were pretty much locked in, the worthiness of the adoption was blindingly obvious at that point, so I can see why tokenmaxxing would be where this ended up, even though (here I disagree with the article a bit) the tools aren't at the "compounding correctness" stage just yet. It's really quite simple: the stick didn't work (telling people in increasingly direct ways to try using AI for stuff), so they tried the carrot.
$deity knows a good chunk of engineers will inadvertently fall for any trick that involves a scoreboard. That holds even when they're perfectly aware they're being tricked.
> If you want to see if the tools work, why don't you just ask your employees? Like any normal employer would?
Again, they did that, they've been doing it continuously over past 3 years. Some people are excited, some people don't care, but some - a population that's definitely overrepresented in HN comments - just stubbornly refuse to try. Now that the answers are in, and they speak in favor of AI, the companies are doing what "any normal employer would": trying to get the stubborn employers to do their job they way their bosses want them to.
(In fact, normal employers would be more eager to fire people who keep refusing top-down instructions - but it's also obvious this technology is experimental; the models and harnesses get more powerful faster than people can learn to use them - so carrots make more sense than sticks in this transition period. Stubborn people begrudgingly using those tools offer an entirely unique perspective and explore use cases and approaches that you won't get from excited adopters.)
Everything is so "blindingly obvious" yet nobody can point to ANY serious peer reviewed studies that prove it.
I'm patient, I'll wait.
Peer review is a technique to get evidence from data when SNR is low. It's not "science", it's just a technique. So is "throwing shit at a wall and seeing what sticks". Don't turn techniques into rituals, and science into religion.
Most of the general LLM discourse in our industry is still closer to "proof of the pudding is in the eating" than to "double-blind studies on large cohorts, p<0.01, effect size is still so small that result is useless in practice"[0].
And we're not talking about curated demos either - most of the contested value can be proven for your own specific cases with little to no expenditure of money and time, at a PoC level (it gets more expensive once you try to operationalize it and find kinks that are hard to iron out).
And that is, the article claims (and I agree), the point of last 6-12 months of tokenmaxxing policies and top-down push - it's putting pressure on people to actually go and do those PoC-s for themselves, because just giving the opportunity and permission turned out insufficient for significant part of the workforce.
--
[0] - Ironically, I remember it was the opposite around the time GPT-4 came out. Back then people talked more about specific claims and demanded measured evidence, because it was hard to get the models to reliably do something interesting. But now that the models can handle bad prompting and can understand you even when you're drunk, suddenly people are denying the general capability of LLMs and asking for randomized control trials.
(For double irony, nowadays one can just ask an LLM for randomized trials; the current SOTA models will happily design you a bespoke eval pipeline if you ask them to.)
FWIW, I think most tokenmaxxing is, to riff of what you said earlier, turning a technique into a ritual and science into religion.
This isn't specific to AI, we've had it before with pretty much everything in software (and since well before software), from "object-oriented solves every problem" to "clean code [where every function is] two, or three, or four lines long", to reporting your daily kloc, to bounties for every bug reported and/or fixed.
Humans do what doomers are afraid AI will do: make a sounds-good utility function (tokens, lines of code, bugs, dead cobras) and get surprised when it is easily gamed for something far less helpful than the vision of whoever set the goal.
You don't need a peer reviewed study to tell you that a heavy rock will fall faster than a light rock.
Which is why we have peer review even for obvious things.
Either I don't understand gravity, or you might want to pick a different analogy...
1. "heavy rock falls faster" is what common sense will tell you (I was literally told this by multiple laypeople just a few days ago when sightseeing atop a tall tower)
2. This is disproven by a trivial experiment that nobody thought worthy of trying for millenia
3. therefore we do need peer reviewed studies to confirm even "obvious" knowledge.
Also, note that GP's parent post about "water being wet" is quite the subject of contention in scientific and philosophical circles, so that wasn't the best example either.
Indeed, as per 2., no one is doing the experiments with rocks of different weight, and sufficient heights to easily measure time of fall. However, people have a lot of everyday experience with feathers, grains, leaves, wood, and rocks, as well as objects of various weight made of metal, paper, plastics. And in everyday experience, the heuristic actually holds out well: lighter stuff falls slower, or gets carried away by the wind.
This "heuristic" is purely empirical. You can't disprove it with peer-reviewed studies, because within its scope, it's literally the most basic, purest form of science: direct observation.
So in 1., the mistake is that of incorrect generalization. "Lighter stuff falls slower" is correct for everyday experience, it's the "therefore, heavy rock falls faster than light rock" is wrong.
Not because it doesn't fall faster, mind you - it does[0] - it's just that everyday experience is dominated by aerodynamic effects, and laypeople sometimes[1] mistakenly assign it to gravity.
Which I guess makes it a great analogy for the LLM story. Turns out everyday experience is actually valid in everyday situations. Generalizing from it is usually badly wrong, even if it sometimes arrives at correct answer for wrong reasons (and at wrong scales).
Generalization is hard.
--
[0] - Surprise. It's actually a heavy idealized particle falls at the same rate as light idealized particle. Actual matter is not an infinitely small point in space, and generates its own gravity field, so the heavy rock will land a tiny bit sooner than the lighter one, because it pulls Earth stronger towards itself - but then only if you drop the test bodies one by one (serially), and not together (in parallel, where the difference cancels out). But then it also turns out the mass canceling out for idealized particles isn't just a mathematical simplification, but a very deep truth about the universe...
[1] - Or don't. The question as phrased is, "does heavy rock fall faster than light rock"? This isn't a "specific physics theory question", it's a "real life" question. Treating a positive answer as belief on gravity is an error made by the asker.
> lighter stuff falls slower, or gets carried away by the wind.
Your examples are of smaller-density or larger-surface-area objects, not lighter ones. A bedsheet is heavier than a penny.
> Actual matter is not an infinitely small point in space, and generates its own gravity field, so the heavy rock will land a tiny bit sooner than the lighter one, because it pulls Earth stronger towards itself
When you're timing how long it takes for the rock to land, you're considering the Earth as fixed and applying gravity to the rock's center of mass. The force applied between the Earth and the rock is F = G * (mEarth * mRock) / r^2
So the force that accelerates a twice-as-heavy rock is twice as large.
But the acceleration of that rock towards the earth is a = F/mRock, so in the end, if the rock is twice as heavy, its acceleration is still exactly the same as the lighter rock's.
> but then only if you drop the test bodies one by one (serially), and not together (in parallel, where the difference cancels out).
What are you talking about!?
If you want to split hairs, you could argue that if you drop them serially you're doing a minute change to the Earth's mass (which is actually so minuscule it makes no difference).
But even in your parallel universe of physics where the "heavier rock pulls the earth towards it", you're reaching a paradox similar to the one Galileo was testing for: if I link the heavy and the light rock together, they should fall slower than the heavy rock alone (because the light rock is slowing it down) but also fall faster than the heavy rock (because the total mass of the system is higher).
https://en.wikipedia.org/wiki/Galileo%27s_Leaning_Tower_of_P...
I am puzzled by the claims upthread though, would love to have their response on it.
I'm afraid I have bad news for you...
I run a small business with two employees.
N=2 here, of course, but one of them will experiment with any new process you introduce (as well as plenty more that you don't!)
The other will keep doing what he's always been doing, even if it's frustrating and inefficient, unless you monitor him and force him to use the new process.
I could imagine most "normal employers" would understand that both type of person exists and, assuming you're getting good first impressions from group A, it's usually better off in the long run to shove the new process down group B's throat.
(This isn't to say that the "Group B" employee is less valuable or anything - he is more conscientious and reliable than anyone else we've ever hired - but just that different people need different management styles)
And your second will be struggling to clean up that mess while also getting their own work done.
Of course, you expect the same level of work from both of them, but because person two has to do a bunch of person one's work as well as their own, person one ends up looking better and gets praised by management.
I'm totally not bitter at all.
> it's usually better off in the long run to shove the new process down group B's throat.
> (…) the "Group B" employee (…) is more conscientious and reliable than anyone else we've ever hired
If employee B is proving themselves to be valuable and reliable, then you should trust them to make the best decisions for how they’re going to go about their work and support them. Leave the door open for them to try different things, but no one likes having processes shoved down their throats (your words). All you’re doing is making them unhappy and more likely to leave to go work for someone who’ll value them like they deserve.
I mean, the difference in the metaphor is that we have pretty fully understood carpentry for many hundreds of years. We still find it difficult to write even simple software to address all our needs, as is evidenced by the insane pay in our industry. Carpenters can suggest tools because they know what's out there. The same was not true about LLMs a year ago.
> That is way too charitable, people were being fired based on these metrics
People get fired for all kinds of reasons including no reason at all. Oftentimes leadership even lies about the real reasons for firing people because they don't sound good!
I'm gonna be blunt: if you're in software and you refuse to use AI for moral reasons, I think you should be fired. There's being principled and there's being obstinate and the difference between the two is how well you can convince people that you _have_ principles. Most LLM-hating people fall short on this point, because
> do I really need to link the Jensen Huang quote?
Sure! Link it again, we all know it's highly immoral when shovel salesmen try to make you want shovels.
> If you want to see if the tools work, why don't you just ask your employees? Like any normal employer would?
I do not like this HN take of "let's do this thing that works great in small companies and then just blindly pretend that it'll also work at the largest companies in the world!" No, this doesn't work at "normal companies" because you cannot "just ask" 30k+ employees what they want.
Employees, like EVERYONE ELSE, are resistant to change. If I, as CEO of a company, want to get my company to try Claude I have to measure tokens to see if it's getting used. That's it. There's no wave of delusion here.
People are stubborn. A lot of productivity improvements had to be almost forced upon farmers, for example. Even when early adopters demonstrated the benefits, a decent fraction of them just didn’t want to change.
This is just a variant of the argument ”people don’t know what’s good for them”. You’re very close to the actual answer, which is that the aforementioned ”manager class” is simply convinced that they understand reality better than those below them, which is quite frankly absurd considering the fact that managers very rarely do any of the ”real work” that these tools supposedly make redundant, and yet they still believe themselves to understand the potential better.
Like when doctors insisted they didn't need to wash their hands (https://en.wikipedia.org/wiki/Ignaz_Semmelweis#Conflict_with...)
or "science advances one funeral at a time" (https://en.wikipedia.org/wiki/Planck%27s_principle)
People are stubborn, but sometimes for good reason. Let the stubborn people hold on to their practices, if the innovators are right they will eventually fold anyway.
Sure. Many have not. I’m thinking of stuff like ox-drawn and then mechanized ploughs, four- versus three-crop rotation, et cetera. The point is there is pushback regardless of benefit and even after it’s been demonstrated. Plenty of people are fine being comfortable. Which is fine. But it also explains why companies and societies with a nudge feature do better.
> if the innovators are right they will eventually fold anyway
Again, sure. If it’s their land, it gets acquired. If it’s your land they’re tilling, you get a say.
I’m not saying all-nor even most—pushback is unfounded. Just that there are plenty of cases where it is, and the solution there is to push through the change.
What happens if employees say no power tools are needed and after a few months a competition shows up with power tools and hires a bunch of noobs and beating your production numbers and sales?
Your employees simply may leave the company and work for them and learn the new culture at this new competitor.
Is there any law which prevents people from moving between companies? No? Then the promoters of that company are going to do what they think is fit to keep them in business and stay competitive. Many times they'll be wrong, sometimes they'll be right.
That did not happened.
So yes, it's a great analogy. We're right now well in the stage where bosses say, "evidence is in and conclusively shows this is useful for us, now the job is figuring out exactly how to work it into our particular business".
This is the most crucial bit. Neither ramming it down developers throats nor rejecting it wholesale is particularly productive. You need the conservative people onboard as well, to discover critical edges and failure modes. Including their criticism in the adoption process instead of bluntly banning it is the smarter move. Of course, there will be a few people who just don't play, they will fold eventually or be let go.
See, table saws are dangerous. Famously so. One of, if not the most dangerous tools available to the general public. They spin quickly with lots of torque and pull things in faster than you can react. Pressure can also send loose pieces of wood backwards at high speed. Fast enough to pass through a person sometimes. It's like being hit with an arrow.
Tablesaw accidents can remove fingers and hands instantly, puncture organs.
They can be used safely but they're circumstantial, the worst thing you can do with a table saw is experiment. Once you realise there's a 12-inch razor sharp blade spinning at 3000rpm with up to 5HP you begin to respect how dangerous it could be and want to warn others.
It's intentionally hyperbolic. But you see what I'm saying here?
And yet, that's very much exactly what happened, well predating the table saw; look into steam engines belt driving saw pit blades for logging.
The table saw itself has evolved in many ways, there's a handheld angle grinder with various blades sub tree.
These attachments: https://www.arbortechtools.com/au/shop-online/power-carving/...
are the literal evolution of wrapping chainsaw about discs on an angle grinder: https://www.afr.com/companies/an-inventor-cuts-in-big-time-p...
It's hard to be hyperbolic about danger and spinning objects, blades, chains after looking into the crazy world of farmers and military civil engineers; a shed built whipper snipper for young trees in a plantation to clean rows made out of chains welded to a tractor rim spinning horizontally hanging down behind rig on small tractor and driven by the PTO is not the scariest thing I've seen.
Long story short, people using tools often make and evolve their own tools to better do their work - and sometimes iterations of those proto-tools are kinda super bloody sketchy. There is care to be taken, there will be close calls.
I know GP said "available to general public", but my mind went straight to PTO after reading "One of, if not the most dangerous tools". Less common to see (especially in cities) than saws, but I think larger proportion of people understand they have to be careful near table saws than PTO shafts.
> after looking into the crazy world of farmers and military civil engineers
The whole history of aviation and space exploration is chock full of engineers, physicists and chemists doing crazy levels of experimentation.
That said, my point was different - unlike GP, I ask to consider workshops that had extensive experience with dangers of powered or high mechanical leverage hardware. It's entirely plausible and reasonable for people running those shops to say, "here is the new dangerous power tool, it's obviously pretty useful (ask your friends at $X or $Y if you don't see it), figure out how and where to best work it into our specific workflows, so it makes us most bang for the buck".
The danger may not be "personal injury before you can react", but it is both parts separately, as there's reports of them giving unsafe advice and also of them performing undesired tasks faster than humans can react.
https://techcrunch.com/2026/02/23/a-meta-ai-security-researc...
For everyday tragedy: https://en.wikipedia.org/wiki/Deaths_linked_to_chatbots#Over...
For mass devastation and warcrimes: https://futurism.com/artificial-intelligence/us-military-elo...
(That said: while I regard "AI doom is marketing hype" to be a conspiracy theory when applied to OpenAI and Anthropic, public statements from this guy are absolutely a case where I'd say hyping up destructive power is the point of his job: https://www.ai.mil/About/Leadership/Bio-Page/Article/3940370...)
In in total agreement with you though, forcing tools on employees is very dumb and is terrible leadership. Ask your people what they need to be optimally exceptional and go get them it. Then let them get on with it.
Some employees want AI tools, others don't. Standardizing SDLC workflows > each person does their own thing. So now you have to choose: do you require AI tool use that fit into a new SDLC? Or don't you?
As long as there's evidence that work meets quality gates for any required customer audits, and your customers are happy and in the loop that AI is a thing that may or may not be used to produce the service, then those engineers that want it can have it and those that don't, don't.
Feels like a revision to an SDLC rather than a new one. Without seeing the SDLC it's hard to find common ground though. It really depends on how it's written and implemented and of course: culture. In the example we're working from sounds like the tools are being forced on people, and that's less infosec, SDLC and more unbearably bad leadership.
This is how it's gone down throughout history. It's why we remember the Luddites, textile workers who started smashing stocking frames and power looms because the machinery was introduced over their objections. The whole goal was to undercut the craftsmen's wages and bargaining power.
So, no, your expectation to be consulted was never going to happen and has not happened throughout history as industrialization has advanced.
You're far too charitable. Understanding has nothing to do with it. Big companies are too far insulated from bad metrics. Middle managers get away with anything and everything because their decisions are too far removed from reality. And they're nowhere to be seen when the other shoe drops. And they'll just leave to a promotion elsewhere if they stay and results are bad.
Everything is far removed from reality in bigco. So you get a bunch of theater and house-playing with "data-driven" posters up on the wall. It's a show that everyone is aware of and seemingly we all still attend.
The mandate was literally “the more sawdust you create the more money you’ll make”. Nothing of value is learned by that mandate. Sure it’ll make people use power tools but it won’t cause anyone to learn how to use them to make furniture.
They might understand the danger of bad metric but that doesn’t mean they aren’t victims of them. If there was intentionality here it was lazy as hell at best.
from my time in FAANG... that seems about correct. Probably the people at the absolute top don't want to just pointlessly burn tokens, but pass that down the chain and eventually the rumor mill turns that into "tokens are an input for your performance review" and people start running Wiggum loops to fix minor typos or linters or something—especially if you do it at a time when every company seems to be doing layoffs.
Or count the fingers, I guess. It's all fun and games until someone looses AI.
They don't. They want some metric to support what they want to do and don't care about good metrics at all.
I've spent the vast majority of my career in FAANGs and it's been the pattern everywhere.
Right now my org has a senior director who is constantly battering managers to tell their reports to fill out the weekly surveys.
Why are the employees not filling out the surveys? Because instead of the old once a year large survey with questions about various levels (including local teams where management cared about the numbers and I could see the actions they took) we now get a survey every week with questions that are meaningless and I have no answer for.
"How does team X deliver on its priorities"?
Team X has O(10K) peoples and a barely countable infinity of projects. Most of which I don't know about and most of which I'm not supposed to know about since things are compartmentalized. So I don't know what team X's priorities are, I don't know how they deliver on them, and I never will know. Asking me and my colleagues is a waste of time and money.
...but none of that matters because the directors want "data" and they want a dashboard showing that we're all giving them "data".
People are (in this analogy) building sawdust farms there.
This, obviously, presumes that the person managing this hypothetical carpentry shop knows what they are doing. It's almost laughable.
In truth the carpentry shop owner manages on vibes, has no idea what employees do and also doen't trust them, and tells employees he wants to see a lot of sawdust in the workshop floor.
This is what's happening here, you have people setting up two chatbots to churn useless tokens at each other, making only sawdust.
I contend that tokens per se are actually a waste product, or at least non-value add. The end user doesn't actually care how many tokens were used to make a thing. If you could get the same result with fewer tokens, that would be an improvement.
But to make this work, you cannot tell your workers that you are looking for sawdust, because you just gave them tools that make sawdust very easily.
Though I understand that gets social validation from other people with no actual experience.
Ugh. Tell me you're early in career without telling me. Sophomoric take.
Have we? Is it generally the case that the more tokens you spend, you better results you get? This take is so weird I suspect author somehow financially benefits from tokenmaxxing.
They might own a chunk of NVDA.
"Most teams haven’t yet figured out how to build their own Ramp Inspect or Stripe Minions (if that’s you, reach out — we can help!) but basically everyone is at least using cursor in the side bar."
Here’s what they said, $$$ aside:
> That’s no longer true. We’ve entered a different regime, where spending more tokens generally results in better results. We call this “compounding correctness” — the more tokens you spend on getting a task correct
> Compounding correctness flips the calculus. If more token spend leads to better outcomes, then you’re going to want to spend a lot of time running tokens. Which sure as hell sounds like tokenmaxxing to me! The original incentives to tokenmax are gone, but eventually folks will realize that a new and more powerful incentive has take its place.
> There were ways to get loops to work, but it was hard. You had to think a lot about how to prompt the agent, which in turn required a pretty deep familiarity with how these things work.
> Now, though, it’s easy. Compounding correctness makes it easy
Go on, tell me I’m quoting the OP out of context.
It’s pretty clear this person believes in compounding correctness, while other, more serious people (1) are perhaps more skeptical.
..and Armins company owns pi. You can’t get much more all in on AI.
Compounding correctness sounds cool, but the real examples of people spending lots of tokens are not compounding correctness; they are wide parallel exploration; like Mythos. The OP is confused, and wrong; they’ve made some basic (flawed) assumptions, and based their entire reasoning on them.
…and are selling AI things. How surprising.
Yes, because they are selling AI services. The article is an ad.
https://www.anthropic.com/engineering/multi-agent-research-s...
Their findings suggest multi-agent systems result in better performance attributed mostly to token usage (80% of variance).
For companies that have measured performance based on token spend, they can now dial it back. Employees have learned to leverage AI for things they wouldn’t have prior. Now they know what’s possible and what’s not.
No one is stupid enough to always measure performance based on token spend and have unlimited budget. It was always a temporary thing to transition the employees to a new world.
Management felt like employees weren't leveraging AI fast enough. That's why in 2025, there were many mainstream articles about how CEOs were forcing their employees to use AI or get fired. Tokenmaxxing was just the other extreme. Companies will arrive at an equilibrium.
There's no need to overthink this.
Edit: One reply cited this X post as an example of why management needed to do this. Trying to change a company with hundreds/thousands/tens of thousands of employees is hard. You have to send one simple message at a time. https://x.com/danluu/status/1487228574608211969?lang=en
Most companies focused entirely on doing "what everyone else is doing" at best or "to see if Programmer Joe can be as productive as the entire team so we can fire the rest".
And many indeed fired employees in droves because they were "underperforming in token spend".
This is true of my current overlords. It slipped recently that the reason they went AI-nuts was that a competitor had announced going “AI first” and the market responded excitedly. Not because they thought it was a good idea: because the market got excited and they didn't want to get left behind.
This is quite a change as our market is financial services and I remember a time when we had to support decades old browsers (one large UK bank who I won't name here had IE6, and only IE6, on many of its user's machines until ~2017) and web servers because they refused to upgrade anything.
> "to see if Programmer Joe can be as productive as the entire team so we can fire the rest"
I'm not sure who Joe is in our outfit, but I'm certainly in the “the rest who are to be fired” set. I've been unhappy in dev & related for years so the AI revolution which I don't care for is where I'm consciously letting myself get left behind to find something else to do with my life. Haven't touched it. Was too late to claim one of the first tranche of Claude licences. And the second. Oops. Maybe I'll use AI in my next big adventure, or maybe my distaste for it all means I have a grand future waiting for me in the hospitality industry!
There's some jobs I'd love to do, but I can't face the bullshit of tertiary education again.
Without some sort of ticket, job choices become more limited?
Yes. Or so I'm told, I've not needed to apply for a job for 26 years…
I have something possible available, though whether it still will be in five months (the earliest I'm likely to leave because of [reasons] and a two-month notice period) is a bit unknown. That five months might be ten as there are other major changes in the company (we were bought a while ago) from which the dust should have settled by Feb, and it makes sense to try to hold out that long to see if I'm still hating things with the same passion at that point.
Without that “something” there are less certain tech based options I could look at, and to be honest I really could do with a proper sabbatical style break. The mortgage is paid, I have savings, and no dependents other than the cats, so I have the luxury of considering that option. And if all else fails I've actually done the arithmetic and I can survive on minimum wage for an extended time if I need to, and hospitality work is something friends can get me into above the many others looking (that bit is less of a joke then people assume: it is seriously part of my plans D & E if I can't stick with A and B & C completely fall through).
1. Source for that 4-word quotation? I googled it, but it appears you are the only person who has ever said it?
2. Even if you made up the quote, source for the claim that "many" "fired employees in droves" for "underperforming in token spend"? (Again, even if the companies never used those words, I'm still interested in the source for the claim about many companies firing employees in droves for low token use.)
I’ve had multiple instances of taking months/years to get some devs to use a more sophisticated git client than GitHub Desktop (so they could properly do anything but the most trivial merges/rebases for example). Or to learn how to use the debugger instead of just printing/logging for debugging. For some of them getting them to seriously figure out how to better use AI required a bunch of repeated prodding.
Funnily enough a few years ago they enthusiastically jumped on copilot’s fancier autocomplete in VS Code, but getting them to really figure out how to get the most out of Claude Code required more pushing.
We are now seeing that Claude Code can do a LOT of heavy lifting in our day-to-day work, but the bulk of our employees are stuck cost-maxing and literally cannot "imagine how you are running into your session limits". "I'm fine with the $20/mo account."
There's a case for the cost-maxing has hurt our company.
Fable, for the few days I had it, would eat through tokens pretty quickly, largely because it tended to work much more on its own. I could give it a task and after asking a few questions it would go off and work for 4-6 hours and be done.
I also run a lot of experiments. I'm trying to be a resource that the rest of my team can learn from as far as what works. For example: when one of the people from our parent company asked about automating payroll entry, I threw their documents and discussion at Claude to see what it'd build. That plus churning on their feedback was ~30 hours of API usage right there.
I'm currently experimenting with "loops", and using codex in those loops to provide feedback and review. That gives me fable-like autonamy (that 30 hours of API usage above), maybe even better. But it uses a lot of tokens. Loops is the bulk of why I got to the weekly limit last week.
Plus I'm having it build an experiment on what my ideal "agent mux" would look like. Herdr is really close, I found it after I started that experiment. Now I'm just letting it run when I have spare usage to see what it comes up with.
There was demonstrably zero cost or consequence analysis, which is also why it was dialed back as soon as the (still) subsidized tokens became just slightly less subsidized, and the wise leaders realized they spent huge sums of money with no way of gauging ROI.
LLMs may have their use cases, but let's not make up free excuses for blithering idiots who, by any rights, should all be fired for cooking up money-burning policies that are textbook implementations of Goodhart's law.
Anyway, just needed to get that off my chest.
Also tokenmaxxing was never an intentional and smart strategy employed by companies like you say. It was a mix of fear of missing out, signaling to investors they were in on the hype and recouping investmenets in data centers
Come on now. Let's not think that we are all smarter than management at these companies.
Your livelihood now depends on tokens remaining subsidized. How long do you think your engineers will continue to have the independent ability to maintain your codebase if the tokens got 20x more expensive?
Buy and sip that intelligence straight from the tap.
Outside of a few well run companies, it's hard not to feel like the average IC is smarter than their leadership.
Big Corporate managers are much more likely to have felt the need to “do AI” from their VPs, who in turn got it from the executive team, who have probably been under fire to produce a coherent magical AI strategy that makes to company scale infinitely while reducing costs. In that environment it’s much more likely to be copy-and-pasted charts from Gartner and buzzwords overheard at conferences, combined with the hope that somebody somewhere will eventually turn it all into something that resembles forward movement.
The presentations were convincing and many bought large upfront wholesale tokens then forced them on their employees.
When the savings didn’t materialize and neither did the profits. And the spend was no longer going down but up, many execs rather than admit they were duped began blaming workers at scale for not making up for their stupid decision.
This is the “get them addicted” part of the platform play.
The problem is unlike any successful drug this was expensive and low payoff.
> Tokenmaxxing was just a way to force employees to start leveraging AI in a meaningful way.
> It was always a temporary thing to transition the employees to a new world.
Trying to understand your justification for rejecting Hanlon’s razor.
Do you have a source for this?
Yes, my own company's decision and logic. the big tech companies needing to pump demand for compute.
Demand is already so large that OpenAI, Anthropic, Meta, Google could not fill it. Tokenmaxxing for these companies strictly to pump fake demand is just plain wrong. The inference demand for these companies internally must be a drop in a bucket in overall inference demand.This reminds me of the popular opinion on HN for return to office mandates as executives wanting to recover their real estate investments.
Also are we sure it's all at arm's length? Barring a full audit, it's not possible to guarantee that there's no round-tripping or overstating of revenue. With Microsoft also being a provider for OpenAI, they could be creatively using set-off, or using SG&A, in order to overstate their revenue/gross margin/inference profit margin. I of course have no proof, extraordinary claims etc. etc. It's unlikely but we should at least debate the possibility. They have such a huge collective incentive to do it.
[0]: https://www.wheresyoured.at/exclusive-openai-financials/ (no affiliation with the website owner, who has a unique bias in this)
No, it was a sinister way to manufacture your consent to cause cognitive atrophy in your employees so that you lose your ability to independently operate your business.
You'll come to realize this once they begin charging you more and more for tokens but you will probably not blame yourself for it.
The argument is tokenmaxxing was put in place by companies, with the goal of causing their employees to lose knowledge?
The whole tokenmaxxing thing started because Jensen Huang said insane things like having a single engineer spend 250k in tokens or he’d fire him; and that OpenClaw was basically AGI.
> No one is stupid enough to always measure performance based on token spend and have unlimited budget.
Yes the people forcing these mandates absolutely are this stupid because that’s what people like Jensen Huang, Peter Steinberger and Boris Cherney were touting. Seriously have you ever actually talked to an average C-Level about AI? They are absolutely cooked.
You’re the one that’s overthinking it.
Or are you just blathering about things you’ve never experienced because you met the “CEO” of a five person company once? I find grand proclamations by people who speak in TikTok absolutely laughable memeing.
(Difference being, one of these groups is just lying about objective reality that's trivial to independently verify, the other one are just unlicensed therapists with thousand years old rituals).
That’s not quite what was said there, he’s budgeting half a devs salary as token spend in a podcast and that if he had a 500k engineer who spent 5k on it at the end of the year he’d go ape shit.
Now you can say that’s wild, and sure, but this is not a standard c suite exec talking it’s the ceo of Nvidia.
Even ignoring other hiring costs this is essentially an argument that Nvidia top engineers should get more than a 50% performance improvement with extremely heavy AI usage. To me, that doesn’t seem like such an enormous statement. For the head of a multi trillion dollar company entirely driven by AI sales arguing it gives a useful benefit to engineers isn’t that odd and betting on a 50% improvement within Nvidia seems kinda normal.
> Seriously have you ever actually talked to an average C-Level about AI?
Yes. Single digit percentage improvements over time would normally excite them, the idea of cappable cost performance improvements that last which your devs actually want to experiment with and a cultural and customer expectation that you’re doing this is pretty enticing. Particularly during a time when tokens were heavily subsidised - isn’t that the perfect time to do it? Now that has ended there’s a huge focus on roi.
Of course not. That is not what it achieved or could possibly achieve.
> Management felt like employees weren't leveraging AI fast enough.
I agree it was about their irrational feelings.
I also agree with the comment you're replying to as well - the vitriol and anger, along with the "this is just another blockchain bubble" type relies is really interesting. It's so surprising to see the variety of (negative) replies and beliefs people have, along with the general distaste/distrust for management. I guess it's also largely a sign of the times since a lot of ICs probably have a ton of anxiety about their career.
This is especially true for the devs who take the code more seriously than the business that employs them. The technical PM who knows a bit of design are suddenly the kings of the company.
Why would those who agree “take the time to reply”? To say what? “This”? “Agreed”? “This guy knows it”? Those comments don’t add anything of value. When you agree, it only makes sense to reply if you have something to say which wasn’t covered by the original argument.
I'm just pointing out that there are equally, if not more, people who agree with me than what the replies seem to suggest.
I agree, but for a completely different reason. A lot of executives simply chase trends. This was another trend they copied from each other. No reason to imagine they carefully studied the issue.
At the IC level, people don't sense the impending urgency for the overall business. They usually sense the urgency for themselves first. AI has completely changed the software industry in 6 months. We went from having AI write some code and copy/pasting to having AI write 99% of the code in 6 months. SaaS went from nice UX and CRUD code logic being a moat to these being nearly free.
Big software companies have to adapt to this new world or they will be outcompeted by smaller, newer, nimbler companies. That's what management is thinking. For ICs, they're usually thinking about their own jobs first.
When everyone was reading about token leaderboards on all of their social media channels (include social news sites like Reddit and Hacker News) it created token anxiety even at companies that didn’t want a leaderboard. Programmers were afraid that their managers would be secretly ranking them based on token usage and they needed to pump up those numbers to avoid layoffs.
Once teams implemented token budgets in response it creates an ugly situation where a few people feel the need to use as many tokens as they can at the beginning of the budget window to stay ahead.
It’s really frustrating to have this phenomenon leak into a company that was never encouraging or looking for high token use.
Instead there was FOMO mass hysteria. Now there is a backlash. And a lot of time and money wasted.
employees who are on the ai bandwagon are there for the free management attention.
Management is cooked because the damn market is hard, money is tight and they can't afford to fight the top down love and $$$ thrown at AI.
If you zoom out, all the real money spent on energy to keep AI alive isn't going to be held in nvidia stock for too long. it will burst, but its stupid to time it.
A sensible organization machinery will move to optimize the metrics that make money. Often times figuring out said machinery takes iterations. Some of them are idiotic (ref: tokenmaxxing) but they are generally directionally correct.
Accenture was.
You spend money for a potential benefit. In this case it’s also a one off cost to find things that can save money over time.
If my productivity is in line with their expectations, I don’t understand why management cares what tools I’m using to do it. No employer ever told me to use emacs instead of vi, even though I’m 10x more productive in one vs the other. So why all of a sudden does management need to micromanage my tools?
Edit: I mean besides the obvious of "because they will fire you if you don't care"
But idk. They're aiming to fire me eventually and have AI do 100% of my job so meh. Fire me now instead of later.
Because doing so increases the value of their stock options. They might privately think it's as dumb as you do, but apparently the stock market disagrees.
Imagine you had a direct report. They were doing just fine, slightly better than a typical report. Then you found out they were writing all their code in notepad - no linting, no automated tests or live updates, no refactoring tools, no highlighting or any code search. They didn’t have any cross code searches and didn’t have any documentation. When they hit a problem, they’d churn away at it and never reach for docs, google or so.
Still, their performance is in line with what you’d expect from someone in their position.
Would getting them to try emacs, vi, linters, etc be micromanaging them? Do you think they’d perform better with them? They are performing in line with expectations for the role, so why bother with something you think would make them more efficient?
I’ve made this obviously over the top, and can hear already replies from other bemoaning my comparison while missing the point — tools do matter and if you genuinely believe that a developer could be more efficient working in a different way it makes sense to not only want them to try it but to actively fund that change. Hell, this is literally what we argue for in training! Spend money to make someone better at their job!
If you think AI tools make you worse or don’t and can’t help, then that’s one thing. But it makes sense for management if they think it might to spend money on it and to get you to try.
Not only this, but wasn’t everyone here shouting about how tokens were subsidised and it couldn’t last? If so, wasn’t the first half of this year a really excellent cheap time to do the maxxing?
Your hypothetical developer wouldn't be using notepad because they're unaware of other editors, they'd be using it because they evaluated other editors and concluded that, for whatever reason, they would be worse for them. I'd be fascinated to hear why they came to that conclusion, but I'm not going to tell them they're wrong if they're performing acceptably, aren't constantly breaking CI because the linter rejects their code, etc. Everyone is different, and I'm not narcissistic enough to think the fact that I would be way less productive without my modal editor, LSP, linter, terminal multiplexer, etc. justifies forcing everyone else has to adopt my exact setup.
Surely for this specific example of managerial stupidity it just is, but I mean more generally, it's a beautiful posting.
I aspire to have this much misplaced belief in any humans at all, let alone CEOs.
"We can't know all the parts of our business that AI can do a good job automating [because it's so new] but we also don't want to be the last to know and outcompeted along the way. Please throw AI at random parts of your job [and we're tracking this] so we can generate feedback from employees on where to invest in additional automation"
My company has since provided a ton of high-value little AI workflows, alongside a handful that didn't pan out. AI-assisted software development is a major change overall, but the general business-process updates from AI are a net-positive to me.
They really don’t IMO. Hell most of the companies pushing these tools don’t even agree what LLMs are for or are capable of. Too many people are trying to use it to cut too many corners on their work (making more work for everyone else) or are using it to attempt things they don’t know how to do, which means they are incapable of vetting the results, (vibe coding anyone?) which means more instances of the first case or even getting hurt.
I wish there was an independent body truly assessing the impact of big tech decisions and running counterfactuals. Instead of accepting nice stories like this as a given.
Using Claude, I recently tried to do something similar for the Covid hiring spree: https://claude.ai/public/artifacts/21bba86a-ad5d-439c-861d-0...
Still leaves huge questions about ROI ($26tln of TAM, anyone???) and doesn't quell the concerns brought forward by AI detractors though.
Pet peeve of mine is nonsensical usage of the x is dead, long live x.
I fear a world where critical software is stood up with increasingly non-human governed abstraction because it [seems like it] works.
Software engineers as the review terminal in a conveyor of business-led code mass production... coming to a company near you?
This is purely for coding and analogues.
Anecdote, I thought so too until the company I work just instated this where you have spend from 35-60K within 6 months. Insanity
In my current company nobody forces you to use more tokens, but you're encouraged to write a 300 lines markdown skill.md which takes 8 minutes and costs 5 bucks to execute. That, instead of writing a 200 lines bash script doing all the same thing, but in a deterministic fashion, completing in under 5 seconds and costing 0 if you're not careful with rounding.
Not necessarily the desired result, but until it's 'done', where the LLM itself is the judge on if the is the case according to the given criteria (often just an updated todo-list). One of those extremely simple 'harnesses' (if you can even call it that) was even named the 'Ralph Wiggum Loop' [1] to allude to the braindead-but-persistent tokenmaxxing it results in.
Otherwise they often do a first pass looks good enough but it doesn't actually work.
If you were tokenmaxxing you would understand.
Really? ~4 years ago our CEO hired a consultant to fly out several times to do team building exercises. We can't afford to do our 3-year server refresh cycle, but the consultant was no problem to pay.
We just recently had branding consultants come in and also spent thousands of dollars (AWS charges) on rebranding all our photos. We operate in a captive market, if you want to operate in our market you are required to subscribe to our service, and if you aren't in our market you can't subscribe. Branding at the end of the day drives 0 sales.
Heck, reminds me of the time a company I was working with hired a new CTO and one of the first things he did was as "server renaming scheme" using obscure (to the US-centric staff) city names from around the world (database servers are Swiss city names, web servers are Denmark, storage is Finland). We went from cattle naming to pet naming, for a CTO that lasted ~6 months.
In my experience company leadership is not quite as thrifty as this article likes to think they are.
Or more accurately, "Because this is good for my career."
consider me officially triggered
I really struggle to imagine how anyone in a corporate environment has managed to never run into obvious examples of waste like you describe (overpaid consultants and mandatory budgets are classic examples). Office Space came out 27 years ago and has a plotline making fun of overpaid "efficiency consultants" whose only job is to tell management to fire people.
The precondition for that is competition. If some company has idiot managers that waste resources on idiotic things, they're supposed to be wiped out by the companies that are actually smart.
Capitalism requires constant evolutionary pressure and a sort of government directed corporation level eugenics program to constantly apply that pressure in order to function properly. Without that, it's just distributed fascism.
I would say tokenmaxxing = spending without limits or care about results (and assuming results). The term as it is right now, at least.
When it comes to "using tokens overall", then open source models change the equation and in that case, we will enter a phase of 'maximizing AI usage...but for near-zero marginal increase in costs with increased usage'. But even then, the convo would shift to platform engineering...which would then ask 'what value are we getting out of this?'
OR - cloud model economics change over time and we use cloud models as happily and cost effectively as we do cloud storage now. But hard to say when that comes.
Open to thoughts, though.
Why do such fever dreams occur at all? Are they getting more prevalent? More damaging? Do they jepaordize the global economy? Should they be regulated in some fashion?
I can't prove my case, but I think it's a symptom of media manipulation/consolidation, the 'fiduciary duty' delusion, and that shareholders can hold the puppet strings tighter than they used to. More and more, they place their sillytown bets and expect the plebs to dance to them.
That comes out to spending $300,000 per user.
The idea of tokenmaxxing reaches different companies in different waves, so it will be discovered in waves and outgrown in waves in companies and industries in their own cycle.
In the long run, tokenmaxxing is like drunken sailor spending. Scaling is almost always about a large component of efficiency, and lighting money on fire in the street can only last so long.
I predict startups will continue to tokenmaxx while 40,000+ person companies will become a little more conservative.
Citation needed
So after I got tired of choosing by hand, and therefore also a bit blindly, I created a small tool that runs locally and analyzes conversations to tell you which skills, MCPs, or other things are always unused.
347 items never used · ~19354 dead tokens/session · ~$25.49/month A lot of ECC that I never used but always loaded.
If anyone's interested, I've put it on GitHub, thousandflowers/skillreaper.
In the early days of LLMs, we saw the classic hype-driven bi-modality of opinions. Folks were in the "fake news, fad" camp, or they were in the "omg, take over the world" camp.
Those of us closer to the space, with the awareness to know that there was some truth (and a lot of misjudgment) to go around, were in the middle of nowhere. When I co-wrote some driver code with Chat GPT, other engineers (and even one of our directors) told me to keep it quiet. At the same time I had directors and VPs asking me how we could accelerate adoption. For a while, I had access to a cheat code just because I had the audacity to not ask for permission. Folks were sure I would get in trouble for spending thousands per month in LLM operation, but a handful came along for the ride, burning tokens like firewood and learning along the way.
Tokenmaxxing is probably coming from at least a few things:
1. A course-correction for the practiced frugality that kept folks from jumping in and just learning at the ragged edge.
2. A willful and deliberate recognition that the best innovations in the later phases of a disruptive introduction often come from sparks of ideation in concentrations of activity. In other words, we don't know where good is, and we need to find it. (Charitable interpretation from the article)
3. Recognition that, even if they don't know why, leaders and product owners will get punished for not jumping in and, because of bullets 1 and 2, won't get punished for trying and missing. Even if they have no idea what they're doing, they're going to fake it until they make it (or slide into another job).
This last set is where the pain lives. An organization with healthy and increasing AI tool usage will see elevated token counts, but so too will one using LLMs to rewrite wikipedia articles without the letter "m" to keep token counts high. These are pathological behaviors brought on by conflated metrics.
We had discussions about this in the early LLM days, where my old team was looking to ship new capabilities for older products. There was a lengthy VP-level discussion about getting to "80% usage" of the new system vs the old. Because the new system was a superset of the old, I eventually said "we can do that immediately, but it's a cost goal, where we're just aiming to make our business more expensive to operate, rather than a value goal for our users". We didn't adopt the target, but folks were understandably frustrated that they didn't have a straightforward way to measure and report progress.
Tokenmaxxing is, inevitably, a conflated goal, but it's what we have right now. Take advantage of the moment, learn, build, and keep an eye on levers for efficiency.