Mistral OCR 4.1
docs.mistral.ai
docs.mistral.ai
Nothing special about this model for overly-detailed work like mine.
It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could perform better. It's worth mentioning that even the highest-end models do a pretty poor job with intricate text like mine.
Mistral's one advantage is that Anthropic now flags OCR, because they don't allow anything that could be considered "reproduction", even of work for which you own the copyright. So my new workflow is Mistral OCR for the actual OCR, followed by a proofreading pass by Claude (which is allowed). Claude is obviously more expensive, but it caught entirely hallucinated sentences created by Mistral OCR 4.0, so I was glad for the backup check.
I feel Anthropic is destroying itself with all these restriction. They got away because their models were the best for coding, but that is not an advantage anymore as OpenAI and other open source are already better.
Also, Suno is so restrictive that it won't make songs of public domain 19th century poems because someone else already made a song and the lyrics are in their song database. And it won't do covers of your own uploaded humming of a novel melody because it matches "humming in an empty room" to some song fingerprint in the database.
They are all trying to avoid these lawsuits by being ridiculously overly cautious. "Your honor, look how much we went above and beyond, even when it made the product worse for legitimate use cases"
I mentioned it in a sibling reply, but here's Anthropic's support document about not using Claude to reproduce content verbatim that already exists, regardless of copyright.
https://privacy.claude.com/en/articles/10023638-why-am-i-rec...
What does this entail? What does Claude do to decide that the text it was provided was hallucinated? Are you telling Claude that the source was OCR'd by another LLM?
I can't speak for Mistral OCR 4.1, but the hallucinations in 4.0 were so egregious (just completely making up new sentences in the middle of a page) that I knew I can't trust Mistral OCR on its own.
It doesn't always get flagged. Single pages are almost always okay. Running a program that sequentially runs single pages through the API is often not okay - I wrote a program in the early 4.x days before the rule came in, that's how I hit it first. But I've also had entire articles go through just fine recently in a Claude Code session (I'd forgotten about Anthropic's rules!), and then others where I get classifier errors by page 4.
The Mistral OCR errors were small in size. Single sentences, formatting errors, paragraphs with newlines. So this was a genuine proofreading job with small changes. For the most part Mistral is actually good, but I can't have it just inventing sentences in the middle of a document. That's where the Claude proofreading pass was most helpful.
I haven't seen any difference in my ocr workflows, what do you mean by this?
I’ve also had it refuse to OCR public-domain books that included content that it didn’t like, such as references to prostitution in 19th-century books about Japan.
I had one session where Claude refused to continue after it hit some kind of guiderail restriction. I couldn’t see what the trigger was, so I started a new session, gave Claude the link to the previous session, and asked it to diagnose the problem. This new Claude said it couldn’t view the exact guardrail issue, but it did suggest a workaround that turned out to be effective.
Thoughtcrime -like territory and self-sensorship. The AI safety lobby is such a vile influence on the freedom of expression and communication via technology (since AI is starting to eat up rest of technology).
I guess the main problem is positioning AI tools as "human-equivalent" creators by the big AI corps. If they were positioned simply as "better OCR and proofreading" people would attribute to them as much responsibility as they would to a - say - typewriter and we would not need to have this nonsense.
I do realize most of the valuation comes from the positioning of "our TAM is the global salary base of 50 trilion and we aim to supesede human workers in the near future" which implies they need to position this technology as "human equivalent" or that valuation is no longer as credible.
At first I thought it was something in the scanned content that was being flagged, but it was the attempt to transcribe that was itself being flagged. Anthropic mention it on their pages:
"Anthropic takes these steps because Claude’s purpose is to generate new content and ideas, not to reproduce content that already exists."
https://privacy.claude.com/en/articles/10023638-why-am-i-rec...
Side note - Claude itself is not aware of this policy, and is unable to see the API responses - the turn just ends. Which turned into a really bizarre failure state where Claude thought I was gaslighting it and kept insisting it could do the work and even had the entire text in memory. Every time it would go to show me and prove it, it would hit API Error 400. I was only able to convince Claude by showing screenshots of my Claude Code screen output so it could see that I was seeing API errors. I've never seen Claude get into that angry & snarky state before, and I hope it doesn't happen again.
That said, models have sometimes surprising weaknesses and a model could be terrible overall but magically work for one type of document.
I haven't been impressed with any of Mistral's models. They obviously realized that they couldn't compete at the frontier so they decided to go for smaller focused models but even those have not been that good.
We moved away from Cursor but I was looking for a model that would help with FIM (fill-in-middle) multiline autocompletion and people were recommending Mistral's Codestral. We gave it a shot and it was lackluster at best.. Even Google's Gemini did a significantly better job than Codestral.
Ultimately Opus-class models got good enough and I don't do much manual coding anymore.
I'd love to see some advancements in traditional OCR based on ideas and concepts we've learned from newer "OCR-like" models since traditional OCR is drastically cheaper.
I really wouldn't know, though. Anthropic models barf out copyright issues for my use case, so I'm unable even to benchmark them. It's a common problem when you're scanning public domain books. Mine are reference texts often cited.
As another user pointed out, it's surprisingly random (task-specific). Llama Scout outperformed Gemini Flash 2.5 on a benchmark I built at the time. I didn't include an OCR models.
Mistral might indeed be the best OCR-specific model for my task, now that you ask. Funny. It's so bad at my work that I didn't register it might be the best in its category. This is just based on vibes from my single scan.
There's always room for improvement, though. I suspect a tool will emerge for highly detailed OCR that implements a nested bounding-box-based multi-scale approach, effectively OCRing small sections at a time and then gradually compiling them by expanding the surface area using the bounding boxes.
I've thought a lot about implementing it anyway.
edit: I see you're asking about the block labels. Leaving the comment in case someone finds it interesting.
The reality is that the US is betting its economy on data centers at great expense and is exposing its economy to great risk.
Also while geographically a lot of the money and processing power is in the US, the US has been relying on immigration to power its universities and especially AI research has roots all over the globe. India, China, Russia, Europe, etc. AI related know how is finding its way back to all these places.
So, I'm not too worried about the long term here. It will be interesting to see if Anthropic and OpenAI survive their IPOs. Seems like a risky financial bet at this point given the apparent lack of a moat. But if it works out, it will result in a lot of that IPO money being invested in data centers abroad. Including in the EU. Because data residency is a thing here and the EU is too big of a market for companies with that kind of valuation to ignore. We also produce a lot of energy infrastructure (e.g. gas and wind turbines). Those data centers will need lots of power.
Demis Hassabis and Google's Deep Mind are all British in terms of location/talent, but they certainly aren't from a coorporate perspective (Deep Mind work might happen in London, but it is owned by Google, as US company)
(I don't know a lot about nuclear/aerospace industries, I'm guessing there more location centered just in virtue of being based around physical things, but I could be wrong)
"Winning the race" doesn't give you that in any meaningful capacity. It gives you, at the absolute most, a temporary window where that's the case. See: nuclear weapons.
Europe already has an innovation problem that's already causing structural economic instabilities which Germany has been struggling (and lately failing) to prop up.
As much as it pains me to say this, AI is already a tech revolution and it seems like Europe is just ignoring it. There's more innovation in 3 blocks in downtown San Francisco than the entire continent of Europe.
QED being first didn't grant exclusivity.
The premise wasn't that AGI isn't useful. There are actually layers to the metaphor where first-mover advantage of AGI is even less meaningful than it was for nuclear weapons, but I leave those as an exercise to the reader to discover.
And I quote:
> Imagine a world where only one country has AGI/ASI.
This is your opinion. Lots of tech money appears to disagree.
You should at least try to explain your contrarian position, or cite your favorite source that makes an argument that you believe.
If AI is different, the onus is on AI folk to explain why, not the reverse.
The frontier labs are racing as hard as they can right now because they know as well as anyone else that they're at best 6 months, probably closer to 3 months, ahead of the generic commodity AI market. The majority of the tech money currently boosting them is a bet that either they can stay that far ahead without tripping, or that a moat can be engineered. It's not at all clear that it can - or, if it can, that it will before the money runs out.
FWIW I explained it fully and provided the argument, along with explicit grounding examples. Can you tell me what exactly you don't understand?
To quote your other post:
> This is the weakest part of the argument. Absolutely no reason to believe the second inventor will be 10x faster.
Nor is there any reason to believe that it will work at any usable speed. An LLM capable of AGI running at 0.000001tok/s isn't a very good head start. Nor is it a good moat if it can't actually realize anything material fast enough. Lord knows America has labor problems.
I'm demonstrating the context is a lot more complicated than "be first". There are innumerable specific examples. They're not guaranteed to happen, that's not the point being made.
Also "no reason" is curious when it's the explicit pattern this exact industry has exhibited. Chinese models lag a few months, and come in swinging with an order of magnitude more efficiency. Does it apply? Who can say. It's certainly on the table. Pretending like it's any more ridiculous than AGI itself is irrational.
> Does this ever happen? Even in traditional manufacturing, the second “inventor” starts behind and has to improve their own process to surpass the first.
To use the industry itself once again: Japan was the first to deep learning. Performance issues prevented them from capitalizing on it. Now they're barely even a player on the field. Interestingly and adjacent to the industry, Japan was second to symbolic AI, and the FGCS was lightyears ahead of America's crufty Lisp ecosystem. For a more modern example, Google was the first to transformers. Much of this revolution is thanks to them. A shame that Gemini is third-rate at best.
It's not unique to this industry. That second inventor more often than not does a better job than the original one. It's a pretty common pattern. Otherwise, we wouldn't need patents.
My "contrarian position" is anything but. There is nothing new under the sun. You don't just need to be first, you need to maintain exclusivity.
This is the weakest part of the argument. Absolutely no reason to believe the second inventor will be 10x faster.
Does this ever happen? Even in traditional manufacturing, the second “inventor” starts behind and has to improve their own process to surpass the first.
We'll see.
You've fallen for the hype. The only way these AI companies can justify their bullshit worth and burn of capital is by promising a magical AGI. It's just marketing and it is still unclear if LLM's will ever be profitable.
If they won't - then EU is doing the right thing laying low.
> Or a world where Europe only gets access to frontier models 6 months later.
What a calamity!
Exactly. People seem to forget that US is burning money at an accelerating rate without any promise of returns. This is a high risk situation. Time will tell.
My own bet is that this won't happen with the current race, but that's another discussion.
Anyways, bit of a doomer take but if I imagine a world with AGI/ASI, countries don't mean shit anymore. Humans are not the top dog anymore.
Who says the LLM race ends at an AGI/ASI finish line in the first place ?
With a lot of Chinese nationals in American companies, and an American corporation owning DeepMind, and a lot of people very upset with all of them at the same time, this is very messy.
Race dynamics increases p(doom) for everyone.
The non-doom scenarios include "utopia for all", and "power flows to investors, not citizens of whichever nation the winning model's corp. was registered in".
Independently, "oh look all the investors went bankrupt" can happen in both "doom" and "normal technology" timelines.
I hope we (EU) don't waste money trying to train local models (which at least some people in Poland try to do), and tries to build our own chips - AI chips have different architecture than regular processor/GPU, and TSMC doesn't need to be winner in this new race.
And if not this, then smaller labs, harnesses and actual application.
Fine tuning an open model to European values is significantly cheaper than making your own model.
Ha. It's a large market of the LLMs consumption. So it which will affect the AI race. Just from other perspective than you assumed.
They can kick into a true war economy and then become significantly more of a threat. They haven't done so yet because the middle-class population likes their luxury lifestyle (for now).
I highly doubt Ukraine will "win" - in fact I believe Russia is ramping up to find an excuse at some point to attack a NATO country, but in a way and during a point in time it's difficult for NATO to retaliate.
Fair enough. But most of the time, the figure is not used to say that Italy is more advanced than Russia, for whatever definition of advanced. It’s a check in the other direction: even if Russia does have advantages, in the grand scheme of things they are not that big.
> They can kick into a true war economy and then become significantly more of a threat. They haven't done so yet because the middle-class population likes their luxury lifestyle (for now).
They did, just in the parts of Russia that are poorer and less-connected (and less powerful politically). They leveraged their huge geographical advantage to spread the pain as much as possible in the regions that don’t matter politically. It is very difficult to see their situation and argue that they are not in a war economy. The middle class in Moscow and St Petersburg are sheltered, true, but still.
> I believe Russia is ramping up to find an excuse at some point to attack a NATO country, but in a way and during a point in time it's difficult for NATO to retaliate.
I don’t think that is likely. But then, of course Trump is making it more and more likely by the day as he insists on getting bogged down in Iran.
I’m glad Mistral is working on useful solutions.
OpenAI/Anthropic is like a retarded little sibling chasing “AGI” and giving up on rich media and other modalities.
OpenAI/Anthropic is the worst of the mainstream AI.
It goes:
1. Gemini
2. Vidu
3. Le Chat (Mistral)
4. DeepAI
5. [insert MiniMax provider]
If you’re interested you can find contact to me via this profile.
3.5 usd/1000 pages is just too expensive…
Google documentai costs 1.5$ per 1000 which is probably better in quality and speed.
I’ve been there, implementing a way to linearise text from a document with pages with 1, 2 or 3 columns, some of them in landscape is a nightmare. And that’s not even considering equations.
In the end it’s way easier to use a specialised model, trained by other people to do exactly what I need.
I thought it's actually a fun problem! Some of the naive solutions (based on alternating horizontal/vertical recursion) are quite elegant and cool.
What was not so fun was to make it work reliably. I ended up with piles of ugly code to handle edge cases, and issues kept piling up. So in the end I was happy to use someone else’s solution.
So, maybe it can be tuned for your usecase but with that kind of investment €3.5 for 1000 pages is a bargain...
There's a big gulf between "it's as if a human being had transcribed it and reconstructed the original document" and "so completely broken it can't be used for anything".
Imagine you have a cheap and 99% accurate ocr. The other 1% you can detect and apply more powerful (more accurate but slower and more expensive) ocr method. What would you use? At scale these things add up.
Because I'm assuming that's why they get to charge more for the right type of customer.
And the deep learning OCR-only models won't censor, but can and do hallucinate. I've yet to see a 'scan with different approaches and reconcile and say you're not sure if they don't agree' system just work for generic complex documents.
I think we also had a layer that for any quote extracted tested it back if it exists within the original.
If you wanted 100% accuracy, I think it wouldn't be too difficult nowadays to guess the font&size&other text settings, and re render the crucial parts.
What's an example of this?
I'm sorry, noob here. I have a special book that I bought which I can open only inside the Kindle app (Windows/mobile). I have been meaning to screenshot the pages and convert them into a document/PDF. What do I have to do to make it fast? Just upload all the screenshots one by one and tell Mistral "Chat" to OCR them?
OCR should be:
1. Privacy-respecting, i.e. running on your own machine without network communications. 2. Fully open-source. 3. Gratis.
The first one is a must, the second is very important for the public interest, and the third one is a nice-to-have.
Mistral does not appear to satisfy even the first-, let alone all three.
Their hosted, API-based service is something like a third of the cost of this model.
The French government is not to be trusted with the data you give it.
https://cybernews.com/security/ants-hack-france-19-million-r... https://www.lemonde.fr/en/pixels/article/2026/08/14/french-t...
Two cyberattacks that we know of happened this year and the extent of the second one is still unclear at this stage.
This is the same government that has pushed for: - Chat Control v1 and v2 (client side scanning) - Id requirement for social media (that just struck down by the courts) - Arrested of the founder of Telegram because he refused to hand over its users data to the French government and used dubious charges to arrest him
After stuff like Chat Control I think they're obviously seeing a big demand for this kind of "internet safety" technology in Europe.