IMO claude, chatgpt/codex, etc should be able to optimize the PDF use case to be extremely token efficient as it's a very obvious use case. But when I start to explain to my wife/friends why it burns through so much quota, I find myself thinking "why should they have to understand this aspect of it". to me, that the details of PDF parsing and extracting are relevant to users (instead of solved such that you don't have to pay attention to it) shows how these tools are not nearly as "ready" as they are made out to be. I may be preaching to the choir on this one, but just my 2c
Source; my last job working with accessibility and that nightmare.
This workflow is highly optimized.
Token economics also are weird. If you design a fancy new frontend that for example uses a cheap model to parse a PDF into text that is fed into an expensive model, you will probably spend more money because you are on API payscale rather than the "max plan" payscale.
For the same reason as why the oil companies want everyone to use large cars.
I'm an engineer and use my coding agent to deal with PDFs all the time. It can reach for unix tools if it needs them.
I don't think I understand why this is a problem - it uses tokens, but it removes drudgery. This is the entire promise of the technology.
This discussion was about measures, goals and incentives. Follow the incentives.
Gemma 4 works perfectly well offline on limited hardware (I have an 8GB video card) and can handle extracting text from image-based PDFs just fine.
Take a PDF -> run it through MarkItDown [1], using the OCR plugin if you need (point it to Gemma 4) -> now you can ask Gemma 4 questions about the (markdown) document.
I am sure Gemma 4 could even create a GUI to make this process very simple for a non technical user.
hell we have restrictive rules for security stuff so in many cases our network engineers are still doing by hand configs for critical systems.
but in terms of token use it's gotta be "take this pdf and parse these 3 columns into 2" or similar
This and replies to this are surreal. It's like everyone simultaneously decided to forget that you don't need claude or whatever to read a PDF. The document is literally made for you to read...
It’s disingenuous to assume every PDF is actually crafted to communicate to its recipients, even more so to pretend LLM users are in a position to understand all the PDFs they receive
There’s a lot of gray area where help understanding a document is fully reasonable
You can rack up token consumption extremely quickly when you embed LLMs into automated processes or products.
I'd be very surprised if these numbers are just typical coding usage with no scripting/pipeline/automation stuff
Using AI to suddenly deliver massive amounts of code without questioning the requirements