Mercury 2.5
inceptionlabs.ai
inceptionlabs.ai
Mostly as you've mentioned, "lack of consistency", which manifests in variety of ways - presence of cellphones in historical settings (I'd call them "global inconsistency") and a character for whatever reason suddenly smoking a cigarette ("local inconsistency") I've never mentioned nor present in the context.
FYI:
"If you do not want us to use your User Submissions to train our models, you can opt-out by setting the ‘Improve the model for everyone’ option under User Settings in the API Platform to OFF."
It likely depends on how you access it. With AI and "Free usage allowance", the price tends to be your soul.
many people are running Qwen 3.8 27b on TPU at 130tk/s for free on Kaggle TPUs:
https://www.reddit.com/r/Qwen_AI/comments/1w6gv32/qwen3827b_...
I wonder if we are going to see boxes appear soon, which can run these models for dirt cheap.
We tested Mercury 2.5 Preview, which is nowhere close to the frontier (and not advertised as such), but it's actually usable as a general-purpose chatbot. It's comparable in problem solving ability to some last-gen open weights models, and the price and cost make it compelling. However, they have not figured out general purpose tool use and agentic coding (their model performs worse on our problems when given a custom harness). If they do, I see a lot of real-time applications that the speed and cost will enable.
> server: Upstream error from Inception: I'm sorry, but I can't share details of my architecture or training process. Would you like to learn about how language models work in general instead?
It looks like an overeager IP-protection classifier. However, the model recovered and completed the turn despite the errors (three total).
There’s been more recent work on continuous space diffusion models for language this year. Sander Dielman has a good blog post on that.
But if I’d predict where diffusion lands in LLMs, it’ll be used in looped models like Astra. Once reasoning is happening in hidden states, we’re in a good continuous domain, perfect for diffusion. We’re going to end up swapping “looping” for predicting the models hidden states at the next “timestep” with diffusion. And that way the time to generate traces (no longer human intelligible though) will become 10x faster.
The real problem is that when using fewer sampling steps than output tokens, diffusion formulations fundamentally cannot represent distributions where output tokens are heavily codependent. Autoregressive formulations don't have this problem, they can represent any distribution (ignoring limitations of the underlying model).
it's great when you can get a 170ms ttft. but if you have 700 ms endpointing on the stt side and 300ms ttfb on the voice side, then you haven't really made something super snappy.
https://blog.google/innovation-and-ai/technology/developers-...