Amazon employees are "tokenmaxxing" due to pressure to use AI tools
arstechnica.com
arstechnica.com
"Research" isn't part of my job title. If you don't know what's possible then why are you deploying it? You should be telling _me_ what's possible. I mean, you _paid_ for it, how can you possibly not know what you were getting?
> in the expectation that you might learn something useful that will be more valuable in the long run.
"I'll take `what even are profits?' for $200, Alex."
An overly generous steelman in my opinion as well. Have 10% of your employees focus on finding ways to properly leverage the new technology - don’t pressure 100% of your employees with bull shit metrics.
Extremely weird take for a knowledge worker.
It's that simple.
(Never mind that these bloggers are just writing ad copy for cloud providers.)
The top-down approach to encouraging (mandating?) AI usage strikes me as infantilizing to the workers, who are perfectly capable of choosing which tools they use and when.
In the early nineties, it was common for experienced electrical engineers to keep on using schematic entry digital design and look down on RTL and synthesis tools, despite that fact the latter was already way more productive. At some point, management had to put their foot down and force everyone to switch to using synthesis.
It's not unreasonable to assume that many people are set in their ways and unwilling to change their behavior without a bit of a push.
If LLMs truly are as good as their proponents say, engineers will use them even if management outright forbade it. The fact that people aren't using them, and have to be forced, is extremely strong evidence that they are not in fact that useful.
Surely you can't argue in good faith today that LLMs aren't useful? It was a valid argument a year ago, but the latest models are absolutely useful at solving whole classes of problems.
They're not perfect, need to be carefully monitored, can cause weird gambling like dopamine rushes and can cause lazy development habits to creep in. But none of those things negate the fact that, in many situations, they are useful.
Short-term, sure in some contexts. Long-term? Nobody knows yet.
Useful = able to be used for practical purposes.
As an extreme example, most effective weapons are useful to the person using them, but aren't necessarily a net benefit.
Is AI going to make life meaningfully better for most people? That's uncertain. Is it useful for the tasks in front of me today? Yes, definitely.
See my other reply in this subthread. For my line of work, they are in fact ridiculously useful.
[1] https://en.wikipedia.org/wiki/Domino_logic
There was no synthesis algorithms that would map VHDL or Verilog designs into domino logic elements at the time. I believe that the most work in the synthesis-to-domino-logic area was done at the beginning of current century.
So, DEC's engineers and, I think, Intel's engineers were doing work using schematics well into 21-st century.
Leaving aside the ethical aspects of using AI (not because they're not valid, because they're off topic for this discussion), in my line of work, the capabilities and productivity improvement of AI are staggering. Most of it is not writing the new code, which is but a small part of chip design, but everything else.
I can't give a concrete work example, but here is an experiment that I ran a month ago. https://tomverbeure.github.io/2026/04/12/AMIQ-License-Key-Ge.... If it can do that, it's not hard to imaging similar use cases related to root causing complex simulation failures. It is frighteningly good at that.
That's a pretty interesting use case. I assume this is for RTL simulation given the thread, but how do you connect the output of the simulator to the AI?
But for large cases, use tools to extract all interfaces from the waveform file and save it as a text file, or add $display statements in the Verilog itself to dump the transactions. A SOTA LLM will eat it up. You point it to the RTL, a log file with hundreds of thousands of lines, and give it a few lines to explain how it is supposed to behave. Just tell it "My simulation is hanging. Figure why." Wait 15 minutes and it will tell you why it hangs and which line to change in your code to fix it.
I've done the experiment after the fact: I had spent ~3 days to fix complicated 3 bugs. I then rolled back the code and told it "Here is the spec. Find all the bugs in this code". It found all 3 bugs in around 30 min. That's when I realized that things won't be the same anymore. (And don't get me wrong: I love debugging simulations.)
>And why not take the alternative approach of identifying the subset of people who have indeed found solid uses and spread their best practices around?
A bottom-up approach has a far better chance of finding those particularly good use cases, and if you lean on the people how found those fits, they're more persuasive than top-down edicts. They actually know what they're talking about. If the point is to leverage AI for better work outcomes, someone with your experience is far more valuable than "here's a dashboard, make the number go up," which seems to be what's going on at Amazon.
I read that BSV source code is about three times shorter than similar design in Verilog and also has three times smaller defect density (defects per significant line of code). So just by changing the HDL from Verilog to BSV one can have nine (9) times less defects in the design.
You include those only in second round along with guidelines and recommendations on how to use it effectively.
Also, we are talking about large companies here. There will be plenty of more suitable seniors.
The new tool might make a lot of that experience obsolete. Also, some people that were good with old tools might be great with the new tool. Some may not.
Overall, I don't think it's a bad idea to burn some tokens (and money) to let people experiment.
You reward me for wasting tokens and punish me for not wasting them, I will maximally waste them and wont "explore hownto make them useful". The latter wastes less tokens and that is punished.
Managing a lot of people at scale is messy and you have to use crude solutions. It's impossible to know everything that's going on.
If you were a manager you wouldn't do any better. Out of the crooked timber of humanity, no straight thing was ever made.
But more specifically ICs tend to want to say "if you just let me do what I know is right it would be fine." That's a trade-off, too, though. That solution means a lot of people will be messing around due to no accountability.
If setting the actual goal you want to achieve as a manager, and then trusting employees to allocate their focus accordingly, you will absolutely have people faffing off (as likely can't be avoided), but at least those who don't will optimize towards what works according to their actual expertise at least.
this means wasting a lot of metaphorical dog food, but now everyone will be 100% how it tastes.
this then allows you to shift client expectations and alter offerings.
...
or it's just dumb mgmt. but let's be charitable.
If you as a “leader” refuse to go along with the crowd and you’re right, then after the dust settles you look like someone who guessed right. Oh and now we’re in a recession so you are probably having a bad time regardless. You maybe get one promotion, congratulations.
If you refuse to go along with the crowd and you’re wrong, you look like a Luddite, you probably got fired at some point along the way and your judgement reputation is hurt.
If you do go along with the crowd and the crowd is wrong, you are just in the same boat as everyone else. You are probably about the same as if you went against the crowd and you were right, possibly even better because it can take awhile to be proven right and you could be hurt in the middle.
So, I think, once something like this picks up enough steam, it’s just logical on a per individual basis for everyone to go along with it, regardless of how they feel about it internally.
But doing so properly requires expending a serious amount of cognitive effort & agile methodology, which is the exact opposite of what Amazon's management has demonstrated here.
It's quite possible they aren't trying to measure performance but are literally just trying to increase token consumption to feed the bubble and hype.
Plus pressure employees may find new unique use cases for AI.
It's like if your goal is inflation, you give out tons of money and as long as its spent, you achieve your goal.
Absurdly wasteful but Goodhart's Law almost never fails.
It makes for pretty charts, extrapolations, and projections.
It doesn’t matter if the numbers are not particularly correct. As long as the data gathering step can be justified it’ll do. Though bonus points if making the number bigger is a good thing (v.s. tracking something like number of sev 1 issues).
> The first step is to measure whatever can be easily measured. This is okay as far as it goes.
> The second step is to disregard that which can't be easily measured or give it an arbitrary quantitative value. This is artificial and misleading.
> The third step is to presume that what can't be measured easily really isn't very important. This is blindness.
> The fourth step is to say that what can't be easily measured really doesn't exist. This is suicide.
— Daniel Yankelovich, "The New Odds"
This why AWS is bleeding good engineers for years. What is left is starting to look like Boeing post McDonnell merger...
They took out a quarter of their documentation page limited real estate, with AI doc shorts nobody asked for, nobody needs, and cant disable.
At Amazon, something like this is likely a closely watched experiment. They knew it would incentivise waste. But they don't know what the other effects will end up being. Nobody knows -- this thread is full of loose speculation. So Amazon runs the experiment and collects the data.
----
The annoying thing about goals and incentives is that they can either be phrased in terms of input metrics (behaviours within our control) of output metrics (the outcomes we want). Input metrics are bad because they lead to skewed incentives and gaming the metrics. Output metrics are bad because they're largely affected by chance and external circumstances. (This indeed means a goal cannot be SMART on its own, because A and R are typically in tension.)
Amazon knows this. Their WBR structure is essentially about trying to set goals and targets for input metrics, and then carefully observing how input metrics correlate with output metrics. They're using a semi-scientific process to tease out the causal structure of their business. I would assume this token target is followed very closely to learn exactly what its effects are on output metrics that drive revenue and cost.
For more on this, I thnk the best public writing is Carr's Working Backwards and Chin has written about it on Commoncog too.
Simpler explanation management has no ideas and goals and this is a replacement strategy. Because they too are affected by "experimental metrics" to a degree, but that doesn't excuse this trite "science".
Any "answer" this would provide wouldn't be of higher quality than this speculation.
My favorite hilarious metric is measuring the amount of work done by counting lines of code written per day
Or by hours spent in the office
https://nordicapis.com/the-bezos-api-mandate-amazons-manifes...
If a manager or a manager's workforce under it just sat around and ignored AI just because it's stupid and irrelevant and useless, they lose one tool to justify their existence amongst their peers who do not express such views. If they sat around and did their jobs as-before WHILE "investing" on tokenmaxxing, they gain a double dip-able vanity metric like "we spent 12.34 quadrillion tokens last quarter" plus "our new method helped us reduce token count by 10^24 this quarter".
You may call it a fraudulent behavior from a hypothetical shareholder's perspective in this hypothetical scenario, which it is, and call it Goodhart's law scenario too, which it also is, but it's a completely normalized behavior in relative terms. Project Hail Mary is a lighthearted work of fiction.
I worked for a healthcare tech startup that made everyone wear fitbits and you got cheaper health insurance premiums if you averaged a higher # of steps every day. People were putting their fitbits on drillbits and whirring them around to log like 20,000 steps a day.
The moment they made it a metric they failed to do anything useful.
Senior management let go our localisation staff. Now they want us to use AI to translate. They still want manual review.
We use Github Copilot at work, we get a measly 300 requests with the budget to go over if necessary. Opus 4.7 or GPT 5.5 would eat all of those up in a day. Are we supposed to be using more than the allotted amount, do management see that as a good thing. Or is it best to stick within the allocated amount. Who knows? Management are playing games everywhere it seems.
One of the weirder things about all this is how arbitrary and non objective the billing structure seems. One of the reasons I'm happy to use it at work, but won't ever personally subscribe. It's so opaque.
Maybe they’re right. But it’s really hard to see how.
I setup entire virtual teams (Dev, QA, product, reviewers etc with the initiating model just acting as the agent manager to keep it's context minimal) to one-shot some stuff and it kept churning and making progress.
Those days are just about over with the change to token pricing but for a time....
Mine are usually giving specification and telling Claude to implement some part of it, referring to existing code base, writing unittests and running e2e until it passes. This can easily take 4-5 hours.
Then again, I have seen colleagues prompt “are you sure?” And other nonsense like that
For speccing things out, I have a back-and-forth with the grill-me skill, break things down into tickets, as well as kicking off subagents. That said, I significantly overestimated the number of human messages I send.
My daily 90th percentile wrt number of prompts sent is sitting at 160 queries / day and average at 97 queries / day.
Ran an analysis of my last 2000 messages, with the following breakdown
Task delegation / execution: 23% Investigation / diagnosis / “what’s going on?”: 21% Planning / architecture / brainstorming: 15% Testing / verification / release ops: 10% Review / cleanup / quality control: 9% Course-correction / constraints / preferences: 8% Agent / ticket / workflow orchestration: 6% Providing context / evidence / pasted material: 4% Social reactions / acknowledgements / vibes: 2% Other: 2%
"You spent $23, over the $20 food limit. Be more careful next time. You spent $600 on tokens, $200 more than the average. Congratulations!"
> whoever spent $600 on Anthropic last night, great job leveraging Al! But to the person who spent $23 on Uber Eats please remember our limit for food is $20 per meal
I can't say that this isn't happening, but at least the parts of the company I get visibility into, what the article describes isn't my experience. There is a lot of interest in using GenAI, but people are mostly getting kudos around creative uses for GenAI, not just for raw amount of tokens. For most scaled GenAI efforts, there is a lot of focus on output metrics (metrics like accuracy, number of findings, number of things fixed, and so on).
I'm surprised how few comments are written with the prior that Amazon managers aren't stupid or uninformed about how incentives work.
My guess would be that someone created the leaderboard without a lot of consultation with managers, and that some employees feel a competitive urge to try to "win" the leaderboard by burning tokens.
LOL, I'd imagine even Amazon HR would be little restraint in showering such praise.
What we can verify is how how Amazon already treats workers, they will surveil anyone within their systems regardless of the futility of said surveillance. Why are we suppose to not believe them using LLM systems as a means to further control their expensive employees from unionizing or seeking out solidarity with fellow workers? All LLMs do is enable tyrannical managers more power to hold over other workers, said workers are forced to engage in self alienation for fear of losing your job or forcing to do meaningless work as that is what's being tracked (and what LLMs excel at producing).
Hardly a good proposition for any worker.
I'm sorry but I fully do not believe you. This is a company that fires workers for taking too long of a bathroom break where said workers piss in bottles for fear of getting fired and you're going "hey guys, it's not too bad. Only some workers get whipped, others don't!"
One of my favorite heuristics/quotes applies here: "no matter how good the strategy, occasionally consider the result."
Want to know if AI is working for your org? Ask yourself/employees to "show me the result." That requires judgment and taste (is the result something of value, or just the appearance of work having been done), but it will also save you a ton of stress and disappointment later.
However I see tons of people on LinkedIn with ways of backing up context, not wanting to lose context, etc.
This seems like another way the system is being misused. Higher context usage also uses more tokens. I suspect you get worse (and slower) output too than a dense detailed context.
a) you find a particular context that executes well and want to preserve parts of it or not have to repeat explanations
b) you want to continue a session so you don't have to rebuild the context from scratch
I think A is something where it's totally reasonable to preserve pieces as part of like a prompt library or equivalent, or directory-specific agent files, that kind of thing.
I think B is much more likely to lead to problems if you do it over a long time, but it can be pretty useful for getting the last drop of juice out of the metaphorical orange.
I think the antipattern (that I've done myself, admittedly) is swapping between different restored contexts for different tasks or roles - at that point you should be either converting it to more durable documentation if warranted, or curating it more specifically than "restore the entire context" even if it's just one-off.
Ideally that replaces the back and forth cycle of it's this, no it's that, it's that for reasons XYZ with a single ingestible blob that gets the agent up to speed.
Sometimes it's better to dump context incrementally, reinitialize the agent with a subset of the context, or manually prime it, then ask it to write documentation as a focused task.
If every exchange is treated as an independent query/response then it's much easier to see how cutting out the fluff using a combination of its summaries and your own helps stay focused.
― Charlie Munger
Hell, throw a Tarot reading in the middle of the loop so the agent has non-deterministic behavior too.
https://github.com/trailofbits/skills/tree/main/plugins/let-...
Amazon management wants to play five-dimensional chess? Play Balatro instead.
Where? What industry, what kind of projects? The only one where I can imagine it to be true is vulnerability research, and I imagine all the low-hanging fruit to be picked soon
It will spin up a boilerplate uboot or BSP config no problem. I still go in and manually check and add peripherals, but opus 4.7 is terrifyingly smart.
Need to modify or add a new peripheral, it's there no problem. Or in a bare metal project, I can point it at an STM32 cubemx starter repo and ask for a feature (set up the ADC on pins 4 and 7, ask me for parameters) and it's just done. I do in a day what would probably take me 2.
It doesn't help me with reviewing others' work, or planning (I maintain that these are manual tasks). So yeah, I agree with the 40-60%. The parts of my job it helps, it really helps.
My experience is it will attempt read from the wrong memory block resulting in garbadge. But that's a while ago so maybe LLMs have gotten better.
I didn't even need that bootloader, just didn't like the fact that Adafruit one takes too much space :)
I'm confused, isn't the whole point of using the STM32CubeIDE that all the peripherals, like say setting up an ADC on pins 4 and 7, are checkbox features?
It also generates a ton of bloat and comments.
We started working on a new product a few months ago and it's really dangerous up front on an empty code base. It can quickly write more code than you can comfortably understand. The more serious danger is when three people are all doing that at once. I had to bring this up at meetings and try to get a better review culture going.
Now that we're a few months in and changes are more targeted additions to an existing system we're happy with, it's _huge_ (which has been my experience on our existing product). I can drop a brief paragraph I speech-to-texted into my agent, give it a general starting place (where I imagine the issue/feature extension point is), and then tell it to do some research and propose a change. I'd guess it's about 50% of the time that I have to update it's implementation plan. Then I let it run (my favorite is setting this up before a meeting) and come back. Then we have to review the code and go from there.
Definitely a 50%+ speed up in some cases, but not all. It's also great for problems that procrastinating, as it reduces friction so much.
At my company(big name, AI beneficiary), middle management seems to mostly be concerned with shuffling chairs on the deck of the Titanic while they wait for their stock to fully vest. There is very little interest in improving anything, just an obsession with risk avoidance and performative sideshows whenever upper management wonders why execution is so poor.
People churning out slop is slowing me down and the full effects of it won't be felt for a while.
Codex was pretty sure something was wrong with the response object being returned by the endpoint in question. It turned out there was a conversion method applied to the endpoint response, which mutated its input. This method had been running w/o problems for a while, until the dev put it in a useEffect. At this point, React dev mode's policy of rendering everything twice kicked in, which caused the second pass through the conversion method to fail on the now-mutated input object.
Codex never even hinted that the conversion method mutating the input could be a problem, nor anything about React dev mode rendering everything twice (specifically to catch problems like this). Apparently, neither of those came up much in its training data.
My point is that this dev seems to have lost, in a few short months of writing everything with Codex, the ability to trace an error from its source (the error trace was being swallowed in a Codex-written catch block that spit out a generic error message). He was completely stuck and just kept doubling down on trying to get Codex to solve the problem, even checking with Copilot as a backup. I'm not optimistic about where this is headed.
In my view you should 1) use AI as a tool to help you learn and 2) write boilerplate you could have easily written yourself. Getting it to think for you is counterproductive (at least until it replaces us entirely).
Everyone I talk to has nowadays KPIs tied to AI usage on their performance evaluation.
It's astonishing how society forgets.
I have an FT subscription and they keep moving toward this kind of narrative first reporting to get clicks. It’s no longer a believable paper.
That said, I’m kind of having a blast using CC in corporate with all the connectors available at our disposal, and I baffled how little some of my coworkers know about what’s available and what the capabilities are. So it’s clear that perhaps some encouragement is prudent for those who are slower to embrace new technologies, but I’m not sure tokencounting and tokenmaxing are the answer.
That may be an enterprise saas is shit problem, but I'm just happy that my employer now has a wiki search that works.
If I do all of this, do I get a promotion?
People use AI differently and they can be equally productive with a variety of token usage quantities.
Also, different kinds of work are differently amenable to using AI.
Using it to grade people is, err, rather unwise.
> That’s my latest joke — that we’ll have to pretend like we used the tools so they can feel validated they’ve spent all this money on hyped up technology. So, yes, it’s em-dashes and “it’s not just this, it’s that …” so they can hopefully leave us alone
Filing JIRA tickets, updates. Opening PRs, having AI review PRs. This will all use tokens.
No need to tokenmaxx, you will end up burning tokens with just regular AI usage
That said, if you can't figure out how to use AI in a software job you should look into it. Not using AI at this point is a lot like not using CAD as an architect.
They also use a bunch of dumb metrics like, total PRs submitted, total comments made on PRs, etc. To the point that, there are multiple heavily used internal tools to game these metrics. Eg, auto-comment LGTM on any approved PR. Thus, making the metrics even worse than they would have been prior.
> Managers are discouraged from using token use to measure performance, according to a person familiar with the matter.
Like CAD and architects, if you're not using LLM's while coding it's an issue, but Amazon is very clear that this isn't an official metric. I would believe managers know how many tokens you're using, but it sounds like they just interviewed a disgruntled employee who didn't like AI and published it.
You're replying to an amazon employee who says they are being used in performance reviews, in comment thread on an article where 2 other Amazon employees say that their token usage is being tracked and they feel pressure to maximize token usage.
Do you have first hand knowledge to refute these 3 people with first hand knowledge?
The CAD thing is incredibly weird. I've never known an architect who had their CAD usage minutes tracked.
Btw I'm a big tech company and I know many people who are "token maxing". It's very common.
Does CAD software regularly generate an incorrect design that results in a catastrophic failure of the building?
No thanks I’ll just watch y’all slip down the slope.
When LLMs are capable of actually doing a good job, then it might be like that. We are not there yet, and we may never be.
AI is genuinely useful for many tasks. But 2x or greater business value from engineering orgs isn’t it. And even if it was business are terrible at measuring value added on an individual basis.
What they can measure though is token use. I’ve heard the same thing from other large companies my friends work for.
It’s bad enough that I’ve moved a significant amount of money out of US large-cap stocks.
You should have asked AI to come up with a better analogy.
Heh. No need to be ashamed, I used to believe them when they lied to me like this too!
"Wow, look at how fast employee # 2 is setting money on fire! Let's promote him!"
Is that in the contract to use AI tools? If not, then what are they on about.
Very very few jobs in the US give you a contract.
Most people look at sea changes come and go. They all have a story of how they "could have bought Bitcoin when it was $100" or whatever. In an org, you don't want to have the story of "we could have done that when nobody else had", so you incentivize adoption of the tool as hard as possible and hope that dipping feet in the water makes people want to swim. If you don't already have a culture of early adoption (and no large company can) then you have to use blunt incentives. I don't think anyone has demonstrated otherwise.
...except each keystroke has an associated cost, the sum of which may equal or exceed my salary.
mass hysteria perhaps?
There used to be a time where people used to die from dancing too much (from my understanding in which hey I can be wrong, I usually am): https://en.wikipedia.org/wiki/Dancing_plague_of_1518
I think that although we wish to consider ourselves as smart and really intelligent but we run on biological machines and clocks which evolutionary have not much of a difference since 1518 or even the times when we used to hunt and forage for that matter.
it's all a political performance.
Remember when they made the decision for return to office and just...decided to post no support data on why they need to do it?
It does not get any better than that
Jensen, Sam, Dario: https://i.imgur.com/AI7rtCY.jpeg
This measuring of tokenmaxxing as a proxy for something beneficial to the company has got to be the single dumbest thing I have ever heard of in my entire software career.
It would be like some company in the dot com era measuring employee's internet download traffic as a proxy for productivity or internet-pilledness.
Why not just reward employees based on who's submit the largest expenses claims? That might have some correlation to work too, right ?!
Hell, I'm in the bowels of Google as an IC and it's hard to understand what adjacent teams are doing. Even harder for management that never gets their hands on anything.
So while you know engineers are probably bullshitting you with fake work, you can at least turn around and tell your supervisor the numbers. It's all a game of plausible deniability.
There should be an anti leaderboard that highlight people under a threshold. Not trying to learn how to use ai while working at a company like Amazon is almost certainly a bad thing, and cause for looking into why.
New hotness: USE AI NO MATTER WHAT AND WE WILL MONITOR EVERYTHING, THOSE WHO REFUSE TO USE AI WILL GET FIRED.
And that's how you slop yourself into (at least) two major downtimes and burn millions upon millions of dollars for zero ROI - but the stonk markets don't care about lost ROI as long as you go along with the AI hype train.