229 karma · joined March 14, 2024
[1]: https://www.mixedbread.com/blog/multimodal-late-interaction-... [2]: https://www.mixedbread.com/blog/wholembed-v3
the thing is most agents waste most of their tokens looking up information which can cause context rot. most small models are not as good as looking up information. this model helps to lookup information for your main agent, which helps you to save tokens and still maintain quality.
> $150M RR on just ads, +3x from August. On <1M users.
now you get why those system are not cheap. keeping indexes fresh, maintaining high quality at large scale and being extremely precise is challenging. by having distributed indexes you are at the mercy of the api providers and i can tell you from previous experience that it won't be 'currently accurate'.
for transparency: i am building a search api, so i am biased. but i also build medical retrieval systems for some time.
we have couple investigative journalists and lawyers using us for a similar usecase.
- the OmniAI benchmark is bad
- Instead check OmniDocBench[1] out
- Mistral OCR is far far behind most Open Source OCR models and even further behind then Gemini
- End to End OCR is still extremely tricky
- composed pipelines work better (layout detection -> reading order -> OCR every element)
- complex table parsing is still extremely difficult
But in general if you don't need to scale crazy Hetzner is amazing, we still have a lot of stuff running on Hetzner but fan out to other services when we need to scale.