HNHacker News
TopNewBestAskShowJobs

breadislove

229 karma · joined March 14, 2024

submissionscomments
breadislove··on Tokens too cheap to meter
One super important thing missing: Speculative decoding. Things like Dflash(2), Dspark etc. help to do one forwards pass and get 6-7 tokens out of it. (For completeness, the embeddings from the forward pass are passed into a diffusion model which predicts the next tokens, and the model just verifies it (very cheap operation)). So we can produce way more tokens for roughly a similar amount of compute.
breadislove··on How good are agents at CAD?
I am curious if there is a way to speed it up. Would be very interesting to know how long a human expert takes to finish the task.
breadislove··on Introducing Toast 1
you can look it up in the blog. RAG is not super great because of two reasons, single embedding vector models are not that good and stopped improving and second most models are not good at looking up information. we spend great time on improving the modeling side by inventing on the indexing level [1, 2]. and now we trained our model to be very good at search. it is matching the quality of Opus 5 and GPT 5.6 Sol while being faster. it helps your main agent to do the task at greater quality, while reducing cost per task.

[1]: https://www.mixedbread.com/blog/multimodal-late-interaction-... [2]: https://www.mixedbread.com/blog/wholembed-v3

breadislove··on Introducing Toast 1
you can try it here: https://dwarkesh-search-demo.vercel.app/

the thing is most agents waste most of their tokens looking up information which can cause context rot. most small models are not as good as looking up information. this model helps to lookup information for your main agent, which helps you to save tokens and still maintain quality.

breadislove··on Introducing Toast 1
hey its a search agent for YOUR own data but can also work over the web. Mixedbread is focusing on providing evidence for agents for your internal data. Toast can interact with any search api. You should be able to provide the SearXNG api to it and it should be good to go. Here the default harness: https://github.com/mixedbread-ai/toast-harness
breadislove··on Introducing Toast 1
you can plugin your existing stack and use it via an openai compatible client.

https://www.mixedbread.com/docs/agent/chat-completions

breadislove··on Introducing Toast 1
yes for the retrieval benchmarks. For officeqa pro v2 we used Codex (as databricks did) and for Harvey LAB we used the vanilla harvey benchmark. For these benchmarks we added minimal tools to use mixedbread search and toast 1.
breadislove··on Introducing Toast 1
Mixedbread Search is a multimodal & multilingual search product, where you can upload any kind of data and make it searchable. Its powered by Wholembed [1] v3, a late interaction retrieval model.

[1]: https://www.mixedbread.com/blog/wholembed-v3

breadislove··on Introducing Toast 1
there is full lore around the naming. i can guarantee you that we are pretty dedicated around our research and product.
breadislove··on Introducing Toast 1
the issue with smaller general models (see at the charts) are way behind the frontier models when it comes to search. we've found that there is huge uplift of having a fast dedicated model. from our perspective, having a very good index is the biggest lever and then having a specialised model.
breadislove··on AMD acquires Taalas to boost inference performance by etching models in silicon
we have not converged at all, if you look at how different the chinese models in terms of architecture you can guess that the labs are experimenting a lot as well. we are seeing all different types of hybrid architectures, different attention methods and so on. Of course on a high level its still a transformer but if you take a proper look we are seeing more divergence then a convergence.
breadislove··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report?
breadislove··on GLM 5.2 is nearly as accurate as a human book keeper
adam, i'd like to get in touch and would love to run the benachmark with mixedbread as a search backend. we are doing this right now with a lot of compliance companies. would be very curious how it improves quality/cost e2e
breadislove··on Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
yes, your are right. what heading would you have taken here?
breadislove··on Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
everything worth writing, you should write yourself
breadislove··on Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
ah whoops, I'll fix it. ty!
breadislove··on Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
The ndcg loss is minimal 90.26 -> 89.65. This means it maintains most of the quality.
breadislove··on Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
to which email did you send it? can u send it to support please?
breadislove··on Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
this is the reason why we report ndcg and not recall. ndcg respects fine grained details so you get the an overview of how much details you are trading off since it would hurt the ranking.
breadislove··on Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
yes exactly.
breadislove··on The Token Compression Illusion: Why I'm Skeptical of RTK
slop complaining about other slop
breadislove··on How we index images for RAG
very bad take. with most modern multomodal models you get way better performance then going to text first
breadislove··on Show HN: Evidex – AI Clinical Search (RAG over PubMed/OpenAlex and SOAP Notes)
this might be interesting: https://www.theinformation.com/articles/chatgpt-doctors-star...

> $150M RR on just ads, +3x from August. On <1M users.

source: https://x.com/ArfurRock/status/1999618200024076620

breadislove··on Show HN: Evidex – AI Clinical Search (RAG over PubMed/OpenAlex and SOAP Notes)
a good system (like openevidence) indexes every paper released and semantic search can incredible helpful since the the search api of all those providers are extremely limited in terms of quality.

now you get why those system are not cheap. keeping indexes fresh, maintaining high quality at large scale and being extremely precise is challenging. by having distributed indexes you are at the mercy of the api providers and i can tell you from previous experience that it won't be 'currently accurate'.

for transparency: i am building a search api, so i am biased. but i also build medical retrieval systems for some time.

breadislove··on Launch HN: Mosaic (YC W25) – Agentic Video Editing
you should check mixedbread out. we support indexing multimodal data and making data ready for ai. we are adding video and audio support by the end of the year. might be interesting for the OP as well.

we have couple investigative journalists and lawyers using us for a similar usecase.

breadislove··on BERT is just a single text diffusion step
or deberta but nevertheless super interesting!
breadislove··on DeepSeek OCR
For everyone wondering how good this and other benchmarks are:

- the OmniAI benchmark is bad

- Instead check OmniDocBench[1] out

- Mistral OCR is far far behind most Open Source OCR models and even further behind then Gemini

- End to End OCR is still extremely tricky

- composed pipelines work better (layout detection -> reading order -> OCR every element)

- complex table parsing is still extremely difficult

[1]: https://github.com/opendatalab/OmniDocBench

breadislove··on Migrating from AWS to Hetzner
we have extremely processing heavy jobs where user upload large collection of files (audios, pdfs, videos etc.) and expect to get fast processing. its just that we need to fan out sometimes, since a lot of our users a sensitive to processing times.
breadislove··on Migrating from AWS to Hetzner
We have extremely processing heavy jobs where user upload large collection of files (PDFs, audios, videos etc.) and expect to get fast processing.
breadislove··on Migrating from AWS to Hetzner
Hetzner is really great until you try to scale with them. We started building our service on top of Hetzner and had couple 100s of VMs running and during peak time we had to scale them to over 1000 VMs. And here couple of problems started, you get pretty often IPs which are black listed, so if you try to connect to services hosted by Google, AWS like S3 etc. you can't reach them. Also at one point there were no VMs available anymore in our region, which caused a lot of issues.

But in general if you don't need to scale crazy Hetzner is amazing, we still have a lot of stuff running on Hetzner but fan out to other services when we need to scale.

Page 1 of 2Next →